FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.
Sources
- T1FACTS Benchmark Suite: Systematically evaluating the factuality of large language modelsGoogle — The Keyword / AI / Research / DeepMind / Developers