How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across…

aiscience

Sources

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure · TechNews