Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
arXiv:2609.28991v1 Announce Type: cross Abstract: Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to…