NVIDIA and Microsoft Unveil 530 Billion Parameter Megatron-Turing Language Model
A joint effort between Microsoft and NVIDIA results in MT-NLG, a 530 billion parameter language model that sets new accuracy standards across numerous natural language processing tasks.
NVIDIA and Microsoft introduce MT-NLG, a 530 billion parameter monolithic transformer language model that stands as the largest and most powerful generative language model trained to date. This massive model contains three times the parameters of the previous largest model of its type and represents a major leap forward in natural language generation technology.
The development team trains MT-NLG using a combination of tensor-slicing and pipeline parallelism across 420 DGX A100 servers equipped with NVIDIA A100 Tensor Core GPUs. This complex infrastructure achieves an iteration time of 44.4 seconds, enabling the successful training of a 105-layer model that utilizes advanced DeepSpeed and Megatron software frameworks.
MT-NLG demonstrates unmatched accuracy across a broad set of natural language tasks, including completion prediction, reading comprehension, commonsense reasoning, and word sense disambiguation. The model sets new top results across several key NLP benchmarks like LAMBADA, PiQA, and HellaSwag, while also showing significant improvements in tasks involving sentence comparisons and relations.