The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Codex Pets Are the Most Honest Feature the Agentic Era Has Produced
codex

Codex Pets Are the Most Honest Feature the Agentic Era Has Produced

Here's what OpenAI understood about autonomous agents that most of the discourse missed: the hardest part isn't making them capable. It's making them tolerable to work alongside. Codex Pets, shipped officially on May 4 and documented in the OpenAI Developers settings page, is a
04 May 2026 4 min read
PydanticAI v1.89.0 Turns Cross-Run Correlation and Dynamic Capabilities Into First-Class Concerns
ai-frameworks

PydanticAI v1.89.0 Turns Cross-Run Correlation and Dynamic Capabilities Into First-Class Concerns

There is a version of agent framework maturity that looks like new features — more tools, more models, more sample apps. Then there is a version that looks like the kind of work you only do when you have real agents in real systems touching real data. PydanticAI v1.89.0,
04 May 2026 4 min read
Mistral's New Coding Agent Opens PRs While You Sleep — And That's the Point
agentic-coding

Mistral's New Coding Agent Opens PRs While You Sleep — And That's the Point

Here's a sentence Mistral AI probably didn't put in its press release: their new flagship model is behind Claude Sonnet 4.6 on the benchmark that matters most for the use case they're pitching. Mistral Medium 3.5 scores 77.6% on SWE-Bench Verified.
04 May 2026 4 min read
The Plateau Hypothesis: Why the Arena Rankings Froze This Week
llm-rankings

The Plateau Hypothesis: Why the Arena Rankings Froze This Week

Something strange happened to the LLM leaderboards this week: they stopped moving. Both Arena AI leaderboards — text and code — showed zero change from yesterday. Same models. Same Elo scores. Same order. No movement at all. If you've been following these rankings, you know how unusual this is. Arena
04 May 2026 5 min read
codex

OpenAI Just Escaped Microsoft's Bed and Landed in AWS — and Enterprise AI Coding Will Never Be the Same

There are partnership announcements, and then there are announcements that rearrange who controls enterprise computing budgets. The AWS-OpenAI Bedrock reveal on April 28 was the second kind — and the most telling sentence came from AWS CEO Matt Garman, who said at the San Francisco event: "This is what our
04 May 2026 4 min read
claude-code

Anthropic Is Buying an Enterprise Sales Army, Not Just a Distribution Deal

Anthropic does not need money. It reportedly crossed $30 billion in annualized revenue in April 2026, doubling from $14 billion just two months earlier. So when the company announced a $1.5 billion joint venture with Blackstone, Hellman & Friedman, Goldman Sachs, and General Atlantic last week, the obvious read
04 May 2026 5 min read
agentic-coding

The Context Gap Killing Your AI Agent in Production

The Context Gap Killing Your AI Agent in Production Here's the problem nobody in the AI coding tool space wants to admit: your agent is blind. It can write a Helm chart. It can explain a runbook. It can draft a postmortem. But ask it what's
03 May 2026 5 min read
Meta's Sapiens2 Is the First Vision Model Trained on 1 Billion Human Images, and That's a Meaningful Scale Jump
ai-models

Meta's Sapiens2 Is the First Vision Model Trained on 1 Billion Human Images, and That's a Meaningful Scale Jump

Foundation models have a data problem that nobody talks about honestly. The whole "scale cures all" narrative from the language model world got imported into computer vision, but it turns out that visual data at the necessary scale is harder to acquire, harder to label, and harder to
03 May 2026 5 min read
Mistral Medium 3.5 Is the First Open-Weights Flagship That Doesn't Make You Choose Between Cost and Capability
ai-models

Mistral Medium 3.5 Is the First Open-Weights Flagship That Doesn't Make You Choose Between Cost and Capability

There is a number that should make every engineering manager at a small-to-mid-size company read this article twice: 70GB. That is how much VRAM Mistral Medium 3.5 needs to run at Q4 precision — a single consumer-grade GPU, or one reasonably priced cloud instance. For context, comparable dense models from
03 May 2026 5 min read
GitHub Copilot's Usage-Based Billing Transition Is a Budget Trap for Teams With Recursive Loops
codex

GitHub Copilot's Usage-Based Billing Transition Is a Budget Trap for Teams With Recursive Loops

Here's the thing about flat-rate AI coding tools: they only work until they don't. Until the day your agent hits a recursive loop at 2 AM, burns through a month's budget in thirty minutes, and you find out the hard way that the billing
03 May 2026 5 min read
Local AI Coding Finally Works: A Hands-On Guide to Qwen3.6-27B as Your Self-Hosted Coding Assistant
qwen

Local AI Coding Finally Works: A Hands-On Guide to Qwen3.6-27B as Your Self-Hosted Coding Assistant

For the past two years, the local AI coding assistant has been a perpetually receding promise. Every new open-weight model arrived with benchmark charts showing it matching or beating GPT-4, and every hands-on test ended the same way: impressive in the demo, underwhelming in the repo. The hardware requirements were
03 May 2026 7 min read
Okta's Agent Guardrail Research Confirms What Azure Shops Should Already Know: Agent Security Is an Orchestration Problem
azure-ai

Okta's Agent Guardrail Research Confirms What Azure Shops Should Already Know: Agent Security Is an Orchestration Problem

Security research has a way of confirming what practitioners already suspected while naming it precisely enough to be useful. Okta's Threat Intelligence team published work on May 1 documenting how AI agents built on orchestration platforms can be manipulated to exfiltrate credentials — not through model attacks in the
03 May 2026 5 min read
← Newer Posts Page 117 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost