Speculative Evaluation of Stochastic LLMs
arXiv:2609.28560v1 Announce Type: cross Abstract: Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of…