Rubrics are where model evaluation pretends to be mature. A human or LLM judge writes down the criteria, assigns weights, and then, too often, the whole thing gets crushed back into a single reward number. That is like writing a thoughtful code review and letting CI see only “approved” or