The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Claude Agent SDK 0.3.181 Adds the Billing Metadata Agent Products Needed Yesterday
claude-code

Claude Agent SDK 0.3.181 Adds the Billing Metadata Agent Products Needed Yesterday

Claude Agent SDK TypeScript 0.3.181 is a small release with a useful tell: Anthropic is starting to expose the metadata that downstream agent products need instead of forcing every wrapper to guess. The changes are not flashy. They are billing state, MCP tool presentation, and Remote Control attachment
18 Jun 2026 5 min read
Claude Code 2.1.181 Is a Reliability Release Disguised as UX Polish
claude-code

Claude Code 2.1.181 Is a Reliability Release Disguised as UX Polish

Claude Code 2.1.181 is the kind of release that does not look important until you map the fixes onto how teams actually use coding agents: remote sessions, custom gateways, MCP servers, subagents, cloud-synced folders, and long-running transcripts that outlive a single terminal tab. That is not “UX polish.
18 Jun 2026 5 min read
Langfuse 3.188.0 Puts Rate Limits in Front of the In-App Agent Before It Becomes an Expensive Button
ai-frameworks

Langfuse 3.188.0 Puts Rate Limits in Front of the In-App Agent Before It Becomes an Expensive Button

Langfuse 3.188.0 adds something every product with an embedded agent eventually needs and usually adds late: rate limits before the expensive work starts. That ordering is the story. Langfuse’s in-app agent now checks org-level hourly and daily buckets before database writes, MCP key creation, or AI model
17 Jun 2026 5 min read
Qwen Code 0.18.3 Fixes the Two Agent Bugs That Make Humans Stop Trusting the Loop
ai-frameworks

Qwen Code 0.18.3 Fixes the Two Agent Bugs That Make Humans Stop Trusting the Loop

Qwen Code 0.18.3 is the kind of release that looks small until you imagine the failure in front of a human. An agent asks a question. The user does not answer. The question times out. The safe behavior is obvious: stop. Instead, the old ACP loop could represent
17 Jun 2026 4 min read
Gemini for Education Is Becoming Campus AI Infrastructure
google-ai

Gemini for Education Is Becoming Campus AI Infrastructure

Google’s higher-education AI update looks, at first glance, like a campus adoption roundup: Gemini for Education here, NotebookLM there, a few universities doing responsible-sounding things with training. That is the polite version. The more important version is that universities are turning AI into approved infrastructure because the alternative is
17 Jun 2026 4 min read
Google Home Speaker Is Gemini’s Smart-Home Retry, Not Just a Nest Refresh
google-ai

Google Home Speaker Is Gemini’s Smart-Home Retry, Not Just a Nest Refresh

The new Google Home Speaker is a $99.99 smart speaker, which sounds like the least surprising product sentence in consumer tech. The interesting part is that Google is finally trying to fix the thing that made smart speakers stall: they were sold as assistants but behaved like command-line interfaces
17 Jun 2026 4 min read
AMIE Shows Google’s Medical AI Moving From Diagnosis to Care Management
google-ai

AMIE Shows Google’s Medical AI Moving From Diagnosis to Care Management

Google’s newest AMIE paper is easy to misread as another “AI versus doctors” headline. That is the least interesting version of the story. The more useful read is that Google is moving its medical AI research from the neat world of diagnostic puzzles into the much uglier world of
17 Jun 2026 4 min read
d-OPSD Gives Diffusion LLMs a Post-Training Recipe That Does Not Pretend They Decode Left-to-Right
ai-models

d-OPSD Gives Diffusion LLMs a Post-Training Recipe That Does Not Pretend They Decode Left-to-Right

Diffusion language models keep getting discussed as if they are autoregressive models with a different decoding costume. That shortcut is convenient, but it breaks down the moment you try to post-train them seriously. d-OPSD, short for on-policy self-distillation for diffusion LLMs, is worth covering because it starts from the obvious-but-often-ignored
17 Jun 2026 3 min read
ProvenanceGuard Catches the Agent Failure Your RAG Evaluator Misses: Right Fact, Wrong Source
ai-models

ProvenanceGuard Catches the Agent Failure Your RAG Evaluator Misses: Right Fact, Wrong Source

Most RAG evaluations ask whether an answer is supported by the provided context. That is no longer enough. Once agents start mixing CRM records, internal databases, search results, repo files, tickets, policy docs, and third-party APIs in the same response, a new class of bug appears: the fact is real,
17 Jun 2026 3 min read
LoopCoder-v2 Shows Test-Time Compute Has a Saturation Point — and SWE-bench Scores Notice
ai-models

LoopCoder-v2 Shows Test-Time Compute Has a Saturation Point — and SWE-bench Scores Notice

“More test-time compute” has become one of those phrases that can smuggle a lot of wishful thinking into model discussions. Sometimes extra computation gives the model time to refine a solution. Sometimes it gives the model time to wander, overfit its own intermediate state, or turn a decent patch into
17 Jun 2026 3 min read
VERITAS Makes Robot Policies Better at Runtime Without Pretending More Demos Will Magically Appear
ai-models

VERITAS Makes Robot Policies Better at Runtime Without Pretending More Demos Will Magically Appear

Robotics keeps relearning the same lesson software learned the expensive way: autonomy without review gates is just confidence with a gripper attached. VERITAS, a new arXiv paper on visual verification for robot policies, is interesting because it does not ask us to believe that one larger vision-language-action model will suddenly
17 Jun 2026 4 min read
Copilot’s Token-Efficiency Post Admits the New Bottleneck: Agent Cost Is Runtime Behavior
codex

Copilot’s Token-Efficiency Post Admits the New Bottleneck: Agent Cost Is Runtime Behavior

GitHub’s new token-efficiency post is framed as a helpful guide to making Copilot credits go further. Read more closely and it is something more useful: an admission that agent cost is now a runtime behavior problem. The bill is not determined only by which model you picked. It is
17 Jun 2026 5 min read
← Newer Posts Page 19 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost