Agent evaluation has been trying to benchmark software systems with model-evaluation tools. That mismatch is starting to break.
The new AgentBeats paper, “Agentifying Agent Assessment for Openness, Standardization, and Reproducibility”, argues for a simple but consequential shift: if the thing being evaluated is an agent, the evaluator should look like