Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

arXiv:2605.04135v3 Announce Type: replace-cross Abstract: LLM evaluations in applied domains tend to reflect models that were already outclassed at time of publication. We observe a publication elicitation gap: the distance between the AI systems generating the results reported in an academic paper…

aiscience

Sources

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation · TechNews