The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
DoorDash’s DSG Paper Makes Search Grounding a Control Plane, Not a Model Checkbox
ai-models

DoorDash’s DSG Paper Makes Search Grounding a Control Plane, Not a Model Checkbox

“The model has search” is a seductive product checkbox and a lousy production architecture. It is fine for demos. It is less fine when you need to know which provider was queried, how many results came back, whether the answer was cached, why latency spiked, whether the output still obeys
18 Jun 2026 4 min read
Rubric-Conditioned Self-Distillation Turns Rubrics Into Training Signal, Not Just Judge Paperwork
ai-models

Rubric-Conditioned Self-Distillation Turns Rubrics Into Training Signal, Not Just Judge Paperwork

Rubrics are where model evaluation pretends to be mature. A human or LLM judge writes down the criteria, assigns weights, and then, too often, the whole thing gets crushed back into a single reward number. That is like writing a thoughtful code review and letting CI see only “approved” or
18 Jun 2026 3 min read
STARE Finds the Token-Level Bug Behind GRPO Entropy Collapse
ai-models

STARE Finds the Token-Level Bug Behind GRPO Entropy Collapse

Entropy collapse is one of those training failures that sounds abstract until it burns a run. The reward curve looks promising, the model starts getting better, and then exploration quietly narrows. The policy becomes more confident before it becomes sufficiently competent. STARE is useful because it treats that failure not
18 Jun 2026 4 min read
RODS Says Tool-Use Agents Need a Curriculum That Moves With the Model
ai-models

RODS Says Tool-Use Agents Need a Curriculum That Moves With the Model

The interesting thing about RODS is not that it generates synthetic tool-use tasks. Everyone is generating synthetic tasks now; some of them are useful, many are benchmark-shaped sawdust. The interesting thing is that RODS tries to answer the question most agent-training pipelines dodge: which synthetic tasks are still worth paying
18 Jun 2026 4 min read
GitHub Issues Gets Duplicate Detection and MCP Field Writes. That Is Triage Automation With Teeth.
codex

GitHub Issues Gets Duplicate Detection and MCP Field Writes. That Is Triage Automation With Teeth.

GitHub’s duplicate issue detection looks like tracker polish. The MCP field-write support is where the teeth are. Together, they point at a more consequential shift: issue triage is becoming structured, agent-operable workflow data instead of a pile of labels and maintainer patience. GitHub shipped a public preview that suggests
18 Jun 2026 5 min read
MAI-Code-1-Flash Is Copilot’s Small-Model Bet on Coding-Agent Economics
codex

MAI-Code-1-Flash Is Copilot’s Small-Model Bet on Coding-Agent Economics

MAI-Code-1-Flash expanding across Copilot is being framed as a model availability update. That is technically true and strategically incomplete. The real product is not the model; it is the router deciding when a small coding model is good enough that GitHub does not need to spend a frontier model on
18 Jun 2026 5 min read
Qwen Code’s Nightly Makes sed -i Accountable
qwen

Qwen Code’s Nightly Makes sed -i Accountable

The least glamorous coding-agent feature is also one of the most important: knowing when the agent changed a file. Qwen Code’s June 18 nightly, v0.18.3-nightly.20260618.bc3e0b405, contains one substantive change, and it lands exactly on that fault line. Supported single-file sed -i substitutions now route through
18 Jun 2026 6 min read
qwen

QwenPaw’s Post-Release Patch Shows What Local Agents Actually Need

QwenPaw’s latest post-release patch is not the kind of release that wins a launch thread. Good. Launch threads are usually where agent products go to cosplay as magic. The useful signal in QwenPaw v1.1.12.post1 is that the team is spending time on the parts of local
18 Jun 2026 5 min read
nvidia

ENPIRE Shows Robot Training Becoming an Agent-Orchestrated Research Loop

ENPIRE looks, at first glance, like another impressive robotics demo: robot arms pushing objects, inserting pins, cutting zip ties, and seating a GPU into a motherboard. That undersells it. The important part is not that a robot learned a task. The important part is that NVIDIA GEAR, Carnegie Mellon, and
18 Jun 2026 4 min read
France’s AI Buildout Is NVIDIA’s Sovereign-AI Playbook With Real Megawatts Behind It
nvidia

France’s AI Buildout Is NVIDIA’s Sovereign-AI Playbook With Real Megawatts Behind It

France’s latest NVIDIA-backed AI buildout is easy to file under “sovereign AI,” which is exactly how a useful infrastructure story gets sanded into policy mush. The more interesting read is that France is assembling the parts of a national AI production system: Blackwell capacity, regional clouds, open-model work, manufacturing
18 Jun 2026 4 min read
Grok Imagine Video 1.5 Fast Is xAI Chasing the Only Benchmark Creators Actually Feel: Waiting
xai

Grok Imagine Video 1.5 Fast Is xAI Chasing the Only Benchmark Creators Actually Feel: Waiting

The interesting claim in Grok Imagine Video 1.5 Fast is not that xAI can generate another short AI video. Everyone with a frontier-media model can generate another short AI video. The interesting claim is that xAI can make a six-second 720p clip arrive in roughly 25 seconds, down from
18 Jun 2026 5 min read
xAI’s Priority Processing Turns Latency Into an Explicit API Purchase
xai

xAI’s Priority Processing Turns Latency Into an Explicit API Purchase

Latency used to be one of those AI-platform problems developers complained about but could not really buy their way out of without a sales call. xAI just made it a request parameter. The company’s new Priority Processing feature lets API users opt into higher scheduling priority by adding service_
18 Jun 2026 5 min read
← Newer Posts Page 17 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost