The agent-security industry has spent too much time asking whether the model said the bad thing. SafeClawBench asks the more useful question: did the agent do the bad thing? That distinction sounds obvious until you inspect most safety evaluations, where a refusal in the final answer can make a system