Function calling benchmarks have spent years grading agents in rooms where the tools behave. ToolBench-X opens the door to the room most teams actually work in: tools drift, fail, wrap fields strangely, disagree with each other, and occasionally make the model look more competent than the system deserves.
The new