Microsoft and NVIDIA Unveil Record-Breaking 530 Billion Parameter Language Model

Microsoft and NVIDIA introduce MT-NLG, a massive 530 billion parameter generative language model that stands as the most powerful monolithic transformer trained to date. The collaboration combines Microsoft's DeepSpeed optimization with NVIDIA's Megatron-LM to achieve this unprecedented scale.

Microsoft and NVIDIA officially unveil the Megatron-Turing Natural Language Generation model (MT-NLG), a massive 530 billion parameter GPT-3-style generative language model. The tech giants claim this new system is the largest and most powerful monolithic transformer language model trained to date, marking a significant milestone in artificial intelligence research.

To achieve this unprecedented scale, the collaborative research team develops an efficient and scalable 3D parallel system. This innovative architecture combines data, pipeline, and tensor-slicing based parallelism to further parallelize and optimize the training of very large AI models without breaking them into smaller, less capable pieces.

The MT-NLG model builds directly upon the existing foundations of both companies, merging Microsoft's DeepSpeed deep learning optimization library with NVIDIA's Megatron-LM. While Microsoft previously uses DeepSpeed to enable the training of 100-billion-parameter models, this new partnership successfully pushes the boundaries of AI capabilities well beyond that previous threshold.

Read More at the original source →