The most honest coding-assistant benchmark is not the one where a model writes a function and leaves. It is the one where the user comes back nine times asking for small refinements until the code breaks in a way that looks reasonable in the chat and wrong in the test