The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe

LiveCodeBench

A collection of 3 posts
CodeChat-Eval Measures the Thing Copilots Actually Do: Break Code on Turn Seven
ai-models

CodeChat-Eval Measures the Thing Copilots Actually Do: Break Code on Turn Seven

The most honest coding-assistant benchmark is not the one where a model writes a function and leaves. It is the one where the user comes back nine times asking for small refinements until the code breaks in a way that looks reasonable in the chat and wrong in the test
25 Jun 2026 4 min read
Bayesian Control for Coding Agents Says Always Run the Tests Is Not a Strategy
ai-models

Bayesian Control for Coding Agents Says Always Run the Tests Is Not a Strategy

“Always run the tests” is good advice for humans and a bad orchestration policy for agents. The human version assumes judgment: run the cheap checks while you work, pay for the expensive suite when the evidence justifies it, stop when another run will not teach you anything. Coding agents often
24 Jun 2026 4 min read
Python-Only Coding Benchmarks Are Lying by Omission
ai-models

Python-Only Coding Benchmarks Are Lying by Omission

Python has been doing too much unpaid PR work for coding models. For the last few years, “best coding model” has usually meant “best at Python-heavy benchmarks, plus vibes from a few repository demos.” That was convenient for leaderboard maintenance, not especially honest about software engineering. Multi-LCB, a new arXiv
19 Jun 2026 4 min read
Page 1 of 1
The LGTM © 2026
  • Sign up
Powered by Ghost