Most coding-agent benchmarks still grade agents like contest submissions: did the patch pass, did the issue close, did the scoreboard increment? That is useful, but it is also a dangerously compressed view of software engineering. RigorBench, a new arXiv paper on “engineering process discipline” in autonomous AI coding agents, argues