Agent reliability usually fails below the layer people are watching. The model gets the blame, the prompt gets rewritten, and somewhere underneath, a cancelled task leaks, a background thread keeps running, a subprocess does not clean up, or a session store writes half of what the rest of the system