Most RAG evaluations ask whether an answer is supported by the provided context. That is no longer enough. Once agents start mixing CRM records, internal databases, search results, repo files, tickets, policy docs, and third-party APIs in the same response, a new class of bug appears: the fact is real,