The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Agent Benchmarks Just Stopped Measuring Models and Started Measuring the Data Center
ai-models

Agent Benchmarks Just Stopped Measuring Models and Started Measuring the Data Center

Agent benchmarks are growing up, which means they are getting less flattering and more useful. NVIDIA’s latest post around Artificial Analysis’ new AA-AgentPerf benchmark looks, at first glance, like the usual hardware victory lap: Blackwell GB300 NVL72 beats older systems by a lot, the chart goes up and to
14 Jun 2026 4 min read
Qwen Code’s Newest Main-Branch Work Is About Surviving Real Agent Operations
qwen

Qwen Code’s Newest Main-Branch Work Is About Surviving Real Agent Operations

Qwen Code’s most interesting work today is not a benchmark, a model card, or another “look, it can edit files” demo. It is a cluster of main-branch patches about how an agent behaves when the world gets annoying: rate limits, dropped streams, browser refreshes, memory pressure, native computer-use drivers,
14 Jun 2026 6 min read
OpenClaw’s Langfuse Session-Grouping PR Gets Observability Right by Refusing to Leak Raw Session Keys
openclaw

OpenClaw’s Langfuse Session-Grouping PR Gets Observability Right by Refusing to Leak Raw Session Keys

Agent observability has a problem that normal web observability mostly avoids: the unit of work is not a request. A useful agent session can span model calls, tool calls, approvals, retries, background continuations, compaction, subagents, and channel replies. If your tracing system can only show isolated spans, you do not
14 Jun 2026 4 min read
openclaw

OpenClaw’s OpenCode Go Context-Window Bug Shows Why Local-Agent Stacks Need Metadata Discipline

Context windows are the easiest model feature to market and one of the easiest runtime features to get wrong. A provider says a model can handle one million tokens. A catalog row says the same thing. The user starts a long coding session expecting the agent to carry a large
14 Jun 2026 4 min read
OpenClaw’s New agents.setDefault RPC Turns Config Surgery Into an Actual Product Surface
openclaw

OpenClaw’s New agents.setDefault RPC Turns Config Surgery Into an Actual Product Surface

The difference between a hackable agent framework and a usable agent platform is often one boring RPC. OpenClaw PR #92957 adds agents.setDefault, a Gateway method that does exactly what it says: make one configured agent the default. That sounds minor until you look at what developers had to do
14 Jun 2026 4 min read
OpenClaw’s Bound-Agent Workspace Bug Is the Multi-Agent Failure Mode Hiding in Plain Sight
openclaw

OpenClaw’s Bound-Agent Workspace Bug Is the Multi-Agent Failure Mode Hiding in Plain Sight

Multi-agent systems do not fail first in the place the demo tells you to look. They fail in the boring glue: which identity got resolved, which workspace got injected, which memory store got consulted, and where the transcript was written after the turn. OpenClaw PR #92961 is a small patch
14 Jun 2026 4 min read
MiniMax M3 on NVIDIA Is a Long-Context Infrastructure Test, Not a Prompt-Box Flex
nvidia

MiniMax M3 on NVIDIA Is a Long-Context Infrastructure Test, Not a Prompt-Box Flex

The least interesting thing about MiniMax M3 is that it has a 1M-token context window. Giant context windows are now the AI industry’s favorite spec-sheet flex: impressive, expensive, and frequently misused. The more interesting story is that NVIDIA is treating MiniMax M3 as an infrastructure problem instead of a
14 Jun 2026 5 min read
Qwen Code 0.18 Turns the CLI Into an Agent Runtime With Desktop, Daemon, Teams, Workflows, and Cron
ai-frameworks

Qwen Code 0.18 Turns the CLI Into an Agent Runtime With Desktop, Daemon, Teams, Workflows, and Cron

Qwen Code 0.18 is what happens when a terminal coding assistant starts admitting it wants to be an agent runtime. The release is too large to read as a normal changelog. Desktop app, daemon mode, ACP bridge, Agent Team, workflow scripts, cron loops, tool-output truncation, telemetry, cross-session rewind, plan-mode
14 Jun 2026 5 min read
OpenCode 1.17.6 Fixes an MCP Honesty Bug: Do Not Advertise What You Cannot Do
ai-frameworks

OpenCode 1.17.6 Fixes an MCP Honesty Bug: Do Not Advertise What You Cannot Do

OpenCode 1.17.6 fixes the kind of bug that should make every MCP client maintainer slightly uncomfortable: the client now says what it can actually do. Not what it might support later. Not what the surrounding ecosystem expects. What it can handle today. That sounds small because capability declarations
14 Jun 2026 5 min read
Codex Alpha 19 Is About Cross-OS Plumbing and Plugin Surface Control, Not Demo Theater
ai-frameworks

Codex Alpha 19 Is About Cross-OS Plumbing and Plugin Surface Control, Not Demo Theater

Codex alpha 19 is not the kind of release that wins a demo video. That is the point. The interesting work in OpenAI’s latest Codex prerelease is not another agent trick; it is the plumbing that decides whether a coding agent can survive real work across Windows, remote execution,
14 Jun 2026 5 min read
OMP 15.12.4 Shows Why Open-Source Coding Agents Are Becoming Provider Control Planes
codex

OMP 15.12.4 Shows Why Open-Source Coding Agents Are Becoming Provider Control Planes

OMP 15.12.4 is what a real multi-provider coding-agent release looks like after the demo ends. There is no single shiny feature that fits neatly into a launch tweet. Instead, the release is full of transport rewrites, retry rules, empty-response handling, OAuth fixes, model registry cleanup, MCP reauthentication, token-accounting
14 Jun 2026 5 min read
CodexBar 0.35 Turns Coding-Agent Quota Anxiety Into a Menu-Bar Product
codex

CodexBar 0.35 Turns Coding-Agent Quota Anxiety Into a Menu-Bar Product

CodexBar is the kind of product category that should not need to exist, which is exactly why it is interesting. A menu-bar meter for AI coding limits sounds like developer desk clutter until you look at the current agent market: Codex has plan-dependent usage windows, Copilot is increasingly tied to
14 Jun 2026 5 min read
← Newer Posts Page 27 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost