PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

arXiv:2609.29504v1 Announce Type: cross Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in…

aiscience

Sources