Rohan Shravan
This systems experience report trains a 120B sparse MoE language model end-to-end on a single 8-GPU node, leveraging reversibility, state-preserving growth, and single-node economics.
Training large language models typically requires hundreds of GPUs, limiting accessibility. Additionally, activation memory and optimizer state memory grow with model size, posing challenges for single-node training.
1) Reversibility: A reversible recurrence stack reconstructs activations during backprop, keeping activation memory flat regardless of model size. 2) State-preserving growth: Each expansion (dense→MoE, shallow→deep, few→many experts) preserves previous weights and is paired with failure cases. 3) Single-node economics: TQP (quantized base expert weights + trained low-rank adapters) reduces optimizer state on expert paths by ~45x.
The 120B model trains on a single 8-GPU node with 8K context, reaching a final training loss of 1.78. Per-domain held-out loss demonstrates targeted capabilities (multilingual Indic, code) are learned by construction. The model family, tokenizer, and training code are released.