The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Claude Fable 5 Tops Arena, But the Fine Print Matters
llm-rankings

Claude Fable 5 Tops Arena, But the Fine Print Matters

The interesting thing about Claude Fable 5 taking the top spot on Arena is not that Anthropic found another few Elo points. That is the normal frontier-model treadmill: publish, benchmark, spike the chart, wait for the next lab to answer. The interesting thing is that the new public #1 model
12 Jun 2026 5 min read
On-Policy Distillation Looks Dense on Paper, but the Parameter Updates Are Sparse Where It Counts
ai-models

On-Policy Distillation Looks Dense on Paper, but the Parameter Updates Are Sparse Where It Counts

On-policy distillation sounds like it should rewrite models broadly. The student generates its own trajectories, the teacher provides dense token-level feedback, and the optimizer gets a lot more signal than sparse reward RL. The surprise in this paper is that the final parameter updates still look selective. “Dense Supervision, Sparse
12 Jun 2026 4 min read
EvoArena Shows Why Agent Memory Needs Git-Style History, Not Just a Bigger Scratchpad
ai-models

EvoArena Shows Why Agent Memory Needs Git-Style History, Not Just a Bigger Scratchpad

Agent memory keeps getting described as a bigger notebook. EvoArena makes the better argument: it should look more like git history. The benchmark suite evaluates agents in environments that change over time: terminal workflows, software repositories, and user preferences. Its paired method, EvoMem, wraps ordinary memory systems with append-only patch
12 Jun 2026 4 min read
RA-RFT Says Retrieval for Reasoning Should Find Analogies, Not Keyword Neighbors
ai-models

RA-RFT Says Retrieval for Reasoning Should Find Analogies, Not Keyword Neighbors

Retrieval-augmented generation has a bad habit: it assumes the useful thing is the thing with the most similar words. RA-RFT is a reminder that reasoning does not work that way. The paper, from researchers affiliated with Meta Superintelligence Labs and Rice University, proposes Retrieval-Augmented Reinforcement Fine-Tuning: a post-training method that
12 Jun 2026 4 min read
Nemotron 3 Ultra Is NVIDIA's Real Answer to Agent Cost: Smaller Active Model, Open Recipes, and Infrastructure Attached
ai-models

Nemotron 3 Ultra Is NVIDIA's Real Answer to Agent Cost: Smaller Active Model, Open Recipes, and Infrastructure Attached

NVIDIA did not just ship another big open model. It shipped a thesis about where agent costs actually go. Nemotron 3 Ultra is a 550B-parameter mixture-of-experts model with 55B active parameters, aimed squarely at long-running agent workflows: coding sessions, research loops, tool-heavy operations work, and multi-agent systems that spend more
12 Jun 2026 4 min read
Copilot SDK 1.0.1 Adds an Experimental API Tripwire. That Is What Agent Platforms Need More Of.
codex

Copilot SDK 1.0.1 Adds an Experimental API Tripwire. That Is What Agent Platforms Need More Of.

The best platform features often look like paperwork. GitHub Copilot SDK v1.0.1 is a good example: the release does not ship a cinematic agent demo, a new chat UI, or a benchmark chart. It ships an experimental API tripwire for Java. That is exactly the kind of boring
12 Jun 2026 4 min read
openclaw

Active Memory Defaults Are Becoming an Agent-Platform Governance Problem, Not a Plugin Checkbox

Memory is where agent platforms stop being chat apps and start becoming governance systems. That is why OpenClaw PR #92253 is more interesting than its diff suggests. The pull request changes Active Memory defaults so the plugin targets configured OpenClaw agents when plugins.entries.active-memory.config.agents is omitted or
11 Jun 2026 3 min read
openclaw

OpenClaw’s Default-Model Doctor PR Is the Gemini Migration Story in Miniature

The ugliest AI-tooling failure is not “the model is unavailable.” It is “the tool says this dead model is your default, then tells you the model does not exist, then fails your agent run anyway.” That is not an error message. That is a trust withdrawal. OpenClaw PR #92292 addresses
11 Jun 2026 3 min read
openclaw

A One-Line Cron Edit Bug Shows Why Agent Orchestration Needs Boring Reliability More Than New Models

The most dangerous scheduler bug is not the one that crashes. It is the one that keeps running, smiles politely, and does the wrong thing in the wrong timezone for weeks. OpenClaw PR #92295 fixes exactly that class of failure. The regression lived in openclaw cron edit <id>
11 Jun 2026 4 min read
openclaw

OpenClaw’s CoreWeave Provider PR Turns Open-Model Inference Into a First-Class Agent Runtime Choice

Provider support usually looks like plumbing until the first time an agent run dies because the model name was right, the base URL was wrong, the context window metadata was missing, and nobody remembered which custom header the hosted inference service wanted. OpenClaw PR #92243 is interesting because it takes
11 Jun 2026 4 min read
Apple Putting Private Cloud Compute on NVIDIA GPUs Is the Privacy Story Builders Should Study
nvidia

Apple Putting Private Cloud Compute on NVIDIA GPUs Is the Privacy Story Builders Should Study

Apple did something unusual this week: it made NVIDIA GPUs part of a privacy story instead of a performance story. That is not the usual role for NVIDIA in AI infrastructure coverage. The default script is simple enough: larger model, faster chip, bigger cluster, lower latency, higher throughput. This announcement
11 Jun 2026 5 min read
Grok Build Gets a Plugin Marketplace, Which Is xAI's Bid for Agent Distribution
xai

Grok Build Gets a Plugin Marketplace, Which Is xAI's Bid for Agent Distribution

xAI just made Grok Build more interesting for the least glamorous reason in developer tooling: distribution. The company has launched a built-in Plugin Marketplace for Grok Build, its terminal coding agent, with an official catalog hosted on GitHub and six launch integrations: MongoDB, Vercel, Sentry, Chrome DevTools, Cloudflare, and Superpowers.
11 Jun 2026 5 min read
← Newer Posts Page 32 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost