Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

arXiv:2609.29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$…

science

Sources