The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
MAI-Code-1-Flash Is Live in GitHub Copilot — And the Benchmark Story Is Only Half the Point
azure-ai

MAI-Code-1-Flash Is Live in GitHub Copilot — And the Benchmark Story Is Only Half the Point

Microsoft's MAI-Code-1-Flash is now live across all GitHub Copilot tiers, confirmed via GitHub Changelog on June 10 with a retroactive publish date that means the rollout has been underway for over a week. The benchmark numbers are the headline material: +16 points on SWE-Bench Pro against Claude Haiku
10 Jun 2026 4 min read
Qwen Code Reverts ACP Memory Commands One Night After Shipping Them
ai-frameworks

Qwen Code Reverts ACP Memory Commands One Night After Shipping Them

On June 7, Qwen Code shipped ACP-mode support for /remember, /forget, and /dream. On June 8, they reverted it. The headline commit in v0.17.1-nightly.20260608.aea34fa2c reads: "Revert 'feat(cli): enable /remember, /forget, /dream in ACP mode (#4811)' (#4818)". That is the most useful
10 Jun 2026 3 min read
Qwen Code 0.18 Preview Fixes the Exact BYOK Failure Mode That Makes Local Agents Feel Flaky
ai-frameworks

Qwen Code 0.18 Preview Fixes the Exact BYOK Failure Mode That Makes Local Agents Feel Flaky

Qwen Code's v0.18.0-preview.1 shipped five commits on June 9, and not one of them is exciting. That is the point. The release fixes three bugs that are collectively a case study in why "bring your own model" is a much harder promise to
10 Jun 2026 4 min read
Codex 0.139 Alpha Is Mostly Boring, Which Is Exactly Why Enterprise Teams Should Read It
ai-frameworks

Codex 0.139 Alpha Is Mostly Boring, Which Is Exactly Why Enterprise Teams Should Read It

The coding-agent race has a visibility problem. While Twitter celebrates demo quality and benchmark scores, the boring infrastructure underneath — the part that determines whether an agent in production actually does what it says it does — moves quietly through alpha releases with release notes nobody reads. OpenAI's rust-v0.139.
10 Jun 2026 4 min read
llm-rankings

The Great Model Split: Arena Says Claude, OpenRouter Says Gemini

The model rankings for June 9, 2026 tell a story that every engineering team should be paying attention to — not because one model won, but because the two major leaderboards are measuring fundamentally different things, and treating them as the same signal is now actively costing teams money. Arena AI&
10 Jun 2026 3 min read
Visual Studio's Copilot Plan Agent Is Microsoft Admitting Vibe Coding Needs a Design Review
codex

Visual Studio's Copilot Plan Agent Is Microsoft Admitting Vibe Coding Needs a Design Review

Microsoft shipped a feature to Visual Studio 2026 that should have arrived two years ago: a Copilot Plan agent that reads your codebase, asks questions, writes a markdown plan, and waits for you to approve it before writing a single line of code. The Plan agent, multi-file summary diffs, context-window
10 Jun 2026 4 min read
Codex 0.137.0 Is Turning App-Server Plumbing Into the Product Surface
codex

Codex 0.137.0 Is Turning App-Server Plumbing Into the Product Surface

OpenAI shipped Codex 0.137.0 with 137 release assets and a release note that reads like a systems programmer's journal: app-server v2 remote-control RPCs, monthly credit-limit visibility, cloud-managed config bundles, machine-readable plugin listing, parallel standalone web searches, hosted image and web tools in more code-mode flows, and
10 Jun 2026 4 min read
codex

Codex Sites Turns the Coding Agent Into an Internal-App Factory — Which Means It Needs Product Governance

OpenAI said more than 5 million people use Codex weekly. The more interesting number in the June 2 announcement is that non-developers now make up about 20% of overall Codex users and are growing more than three times faster than developers. That is not a user acquisition story. That is
10 Jun 2026 4 min read
claude-code

Anthropic SDK 0.102/0.107 Turns Agent Operations Into Typed API Surface

There's a class of SDK bug that doesn't crash anything, doesn't log an error, and produces a failure that looks like your cloud provider is having a bad day. That's the signing-order bug that Anthropic fixed in TypeScript SDK 0.102.0,
10 Jun 2026 3 min read
claude-code

Anthropic Python SDK 0.107.1 Fixes Foundry Auth — The Boring Patch That Decides Enterprise Claude Deployments

There's a particular kind of bug that only appears in production: the kind where two correct things collide. That's what happened with anthropic-sdk-python v0.107.0. The maintainers had just hardened the Foundry auth path to prevent a stray ANTHROPIC_API_KEY from leaking into Microsoft
10 Jun 2026 5 min read
OpenClaw's Gemini CLI Vision Fix Is a Small Shim With a Big Interop Lesson
openclaw

OpenClaw's Gemini CLI Vision Fix Is a Small Shim With a Big Interop Lesson

OpenClaw's Gemini CLI Vision Fix Is a Small Shim With a Big Interop Lesson The bug report in GitHub issue #91739 is 11 lines long and contains more diagnostic precision than most support tickets ten times its length. The reporter describes an environment, a model request, a failure
09 Jun 2026 3 min read
OpenClaw's Plugin EOVERRIDE Bug Is the Kind of Boring Dependency Failure That Takes Agents Offline
openclaw

OpenClaw's Plugin EOVERRIDE Bug Is the Kind of Boring Dependency Failure That Takes Agents Offline

OpenClaw's Plugin EOVERRIDE Bug Is the Kind of Boring Dependency Failure That Takes Agents Offline There is a category of bug that never makes the release notes but costs operations teams hours: the kind where a dependency management change in version N+1 makes every plugin install or
09 Jun 2026 3 min read
← Newer Posts Page 36 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost