Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs
The MaxText team successfully reproduced AI2’s OLMo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original PyTorch-on-GPU reference across pre-training and mid-training stages on all held-out evaluations. The implementation achieved up to 57.4% Model…
Sources
- T1Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUsGoogle — The Keyword / AI / Research / DeepMind / Developers