Pretraining Large Language Models with NVFP4

/ Authors

Nvidia, Felix Abecassis, Anjulie S. Agrusa, Dong Ahn, Jonah Alben, Stefania Alborghetti, M. Andersch, Sivakumar Arayandi, Alexis Bjorlin, Aaron Blakeman

and 80 more authors

Evan Briones, Ian Buck, Bryan Catanzaro, Muya Chang, Jinhang Choi, Mike Chrzanowski, Eric Chung, Victor Cui, Steve Dai, B. Rouhani, Carlo del Mundo, Deena Donia, Burc Eryilmaz, Henry Estela, Abhinav Goel, O. Goncharov, Yugi Guvvala, Robert Hesse, Russell J. Hewett, Herbert Hum, U. Kapasi, Brucek Khailany, Mikail Khona, Nick Knight, Alex Kondratenko, Ronny Krashinsky, Ben Lanir, Simon Layton, M. Lightstone, D. Lo, P. Micikevicius, Asit Mishra, Tim Moon, Deepak Narayanan, Chao Ni, Abhijit Paithankar, Satish Pasumarthi, Ankit B. Patel, M. Patwary, A. Poojary, G. Prasad, Sweta Priyadarshi, Yigong Qin, Xiao-Shuai Ren, O. Rybakov, Charbel Sakr, S. Satheesh, Stas Sergienko, Pasha Shamis, Kirthi Shankar, Nishant Sharma, M. Shoeybi, Michael Y. Siu, Misha Smelyanskiy, Darko Stosic, Dusan Stosic, Bor-Yiing Su, Frank Sun, Nima Tajbakhsh, S. Thomas, Przemek Tredak, Evgeny Tsykunov, Gandhimathi Vaithilingam, Aditya Vavre, Rangharajan Venkatesan, R. Waleffe, Qiyu Wan, Hexin Wang, Mengdi Wang, Lizzie Wei, Hao Wu, Evan Wu, Keith Wyss, Ning Xu, Jinze Xue, Charlene Yang, Yujia Zhai, Ruoxi Zhang, Jingyang Zhu, Zhongbo Zhu

/ Abstract

Large Language Models (LLMs) today are powerful problem solvers across many domains, and they continue to get stronger as they scale in model size, training set size, and training set quality, as shown by extensive research and experimentation across the industry. Training a frontier model today requires on the order of tens to hundreds of yottaflops, which is a massive investment of time, compute, and energy. Improving pretraining efficiency is therefore essential to enable the next generation of even more capable LLMs. While 8-bit floating point (FP8) training is now widely adopted, transitioning to even narrower precision, such as 4-bit floating point (FP4), could unlock additional improvements in computational speed and resource utilization. However, quantization at this level poses challenges to training stability, convergence, and implementation, notably for large-scale models trained on long token horizons. In this study, we introduce a novel approach for stable and accurate training of large language models (LLMs) using the NVFP4 format. Our method integrates Random Hadamard transforms (RHT) to bound block-level outliers, employs a two-dimensional quantization scheme for consistent representations across both the forward and backward passes, utilizes stochastic rounding for unbiased gradient estimation, and incorporates selective high-precision layers. We validate our approach by training a 12-billion-parameter model on 10 trillion tokens -- the longest publicly documented training run in 4-bit precision to date. Our results show that the model trained with our NVFP4-based pretraining technique achieves training loss and downstream task accuracies comparable to an FP8 baseline. These findings highlight that NVFP4, when combined with our training approach, represents a major step forward in narrow-precision LLM training algorithms.

Journal: ArXiv

DOI: 10.48550/arXiv.2509.25149

Pretraining Large Language Models with NVFP4

/ Authors

Nvidia, Felix Abecassis, Anjulie S. Agrusa, Dong Ahn, Jonah Alben, Stefania Alborghetti, M. Andersch, Sivakumar Arayandi, Alexis Bjorlin, Aaron Blakeman

and 80 more authors

/ Abstract

Journal: ArXiv

DOI: 10.48550/arXiv.2509.25149