Most coding-agent benchmarks still have the same quiet flaw: they are too tidy. They ask models to solve tasks that look like engineering work, but often strip away the parts that make engineering work irritating — half-stated user intent, local files, environment assumptions, command output, artifacts, and the long tail of