Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
arXiv:2609.25848v1 Announce Type: new Abstract: A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter…