The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
CrewAI 1.14.8a Pushes Flows Toward JSON-First Runtime Definitions
ai-frameworks

CrewAI 1.14.8a Pushes Flows Toward JSON-First Runtime Definitions

CrewAI 1.14.8a is an alpha release, so nobody should pretend it is a production migration memo. But it is worth covering because it points at one of the defining fights in agent frameworks: should orchestration live as Python code, or should more of it become a declarative artifact
18 Jun 2026 6 min read
Codex 0.141 Turns Remote Agents Into a Secure Execution Fabric, Not Just a CLI Session
ai-frameworks

Codex 0.141 Turns Remote Agents Into a Secure Execution Fabric, Not Just a CLI Session

Codex 0.141.0 looks like a CLI release if you skim the changelog. It is not. This is the release where OpenAI’s coding agent starts looking less like a terminal assistant and more like a secure execution fabric: remote executors, encrypted relay channels, plugin-scoped MCP servers, cross-platform path
18 Jun 2026 5 min read
Microsoft Agent Framework 1.9 Puts Approval Gates and MCP Sampling Limits Into the Harness
ai-frameworks

Microsoft Agent Framework 1.9 Puts Approval Gates and MCP Sampling Limits Into the Harness

Microsoft Agent Framework Python 1.9.0 is the kind of release the agent ecosystem needs more of: fewer “look, it can call tools” demos, more policy wired into the place where tools are actually called. The headline is governance. MCP sampling is denied by default unless an approval callback
18 Jun 2026 5 min read
Gemini CLI 0.47 Ships the Antigravity Migration UX Google Needed Before the June 18 Cutoff
agentic-coding

Gemini CLI 0.47 Ships the Antigravity Migration UX Google Needed Before the June 18 Cutoff

Gemini CLI 0.47.0 is a migration release wearing a routine changelog. That is usually where the important work hides. The day users search for “Gemini CLI June 18” or “Antigravity migration,” they do not need a brand narrative. They need the tool in front of them to say,
18 Jun 2026 4 min read
Splash Canvas Shows Why Creative AI Needs Better Interfaces, Not Bigger Prompt Boxes
google-ai

Splash Canvas Shows Why Creative AI Needs Better Interfaces, Not Bigger Prompt Boxes

Most creative AI products still treat imagination like a support ticket: type a prompt, wait for output, negotiate with the slot machine. Google Arts & Culture’s new Splash Canvas is interesting because it refuses that default. It is small, silly, browser-native, and full of chatty sea creatures — which is
18 Jun 2026 5 min read
llm-rankings

The LLM Leaderboard Split Is Now Impossible to Ignore

The useful LLM ranking story this week is not that GPT-5.5 moved one slot above Grok. That is leaderboard static dressed up as motion. The more important signal is that quality rankings and usage rankings are now telling two different stories: Anthropic owns the top of Arena’s text
18 Jun 2026 4 min read
Diffusion LLMs May Have Found Their Best First Job: Fixing Formal Proofs
ai-models

Diffusion LLMs May Have Found Their Best First Job: Fixing Formal Proofs

Diffusion language models have spent the last year looking for a job description better than “what if chat completion, but stranger?” Diffusion-Proof gives them one that actually fits: formal proof repair. Not general conversation. Not replacing autoregressive models across the board. A narrow, verifier-guided domain where bidirectional infilling is not
18 Jun 2026 4 min read
Agent Security Benchmarks Need to Measure Harm, Not Just Whether the Model Said the Bad Thing
ai-models

Agent Security Benchmarks Need to Measure Harm, Not Just Whether the Model Said the Bad Thing

The agent-security industry has spent too much time asking whether the model said the bad thing. SafeClawBench asks the more useful question: did the agent do the bad thing? That distinction sounds obvious until you inspect most safety evaluations, where a refusal in the final answer can make a system
18 Jun 2026 4 min read
Robot Foundation Models Can Manipulate Objects and Still Forget What Those Objects Mean
ai-models

Robot Foundation Models Can Manipulate Objects and Still Forget What Those Objects Mean

The phrase “robot foundation model” carries an implied promise: if the model learned enough from the web-scale visual-language world, then fine-tuning it into a policy should give the robot useful common sense with hands attached. Act2Answer is a useful bucket of cold water on that assumption. It asks a simple
18 Jun 2026 4 min read
SQL Agents Are Finally Being Judged Like Systems, Not Chatbots With SELECT Privileges
ai-models

SQL Agents Are Finally Being Judged Like Systems, Not Chatbots With SELECT Privileges

The interesting thing about SQL agents is not that they can write a cleaner SELECT. We solved the “generate plausible SQL from a natural-language question” demo years ago, then spent the next few years discovering that plausible SQL is not the same thing as data work. Real data work is
18 Jun 2026 4 min read
GitHub CLI 2.95 Gives Agents a Safer Way to Read Repos Without Cloning the World
codex

GitHub CLI 2.95 Gives Agents a Safer Way to Read Repos Without Cloning the World

The most important coding-agent feature GitHub shipped this week is not another chat panel. It is a pair of preview CLI commands that read exactly one remote file or one remote directory without cloning the entire repository. That sounds small because we have trained ourselves to confuse larger context with
18 Jun 2026 5 min read
GitHub’s Release Notes Now Credit the Human Behind Copilot PRs — Small UX, Big Accountability Signal
codex

GitHub’s Release Notes Now Credit the Human Behind Copilot PRs — Small UX, Big Accountability Signal

Attribution is the kind of product detail nobody cares about until it becomes the only thing that matters. GitHub’s latest generated-release-notes change looks tiny: when Copilot cloud agent opens a pull request, release notes now credit the human who asked for the work alongside @copilot. But small metadata changes
18 Jun 2026 5 min read
← Newer Posts Page 18 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost