Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
arXiv:2608.18744v2 Announce Type: replace Abstract: Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an…