arXiv
Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining
arXiv:2609.25482v1 Announce Type: cross Abstract: Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the…