Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction
arXiv:2609.26268v1 Announce Type: new Abstract: A next-token probability says what a model predicts, not how much training support lies behind it. A Dirichlet head can represent this distinction by separating mean $m$ from concentration $S$, but decoupling does not identify what $S$ means. Here we…