Microsoft and NVIDIA Unveil 530-Billion Parameter Megatron-Turing Language Model

Microsoft and NVIDIA collaborate to build MT-NLG, a 530-billion parameter generative language model that sets new accuracy standards across various natural language tasks.

Microsoft and NVIDIA introduce the Megatron-Turing Natural Language Generation model (MT-NLG), a massive 530-billion parameter transformer language model that stands as the largest monolithic model trained to date. This breakthrough results from a collaborative effort to push the boundaries of AI scale by leveraging advanced parallelization techniques.

The model relies on a powerful combination of Microsoft's DeepSpeed and NVIDIA's Megatron-LM technologies to handle the immense computational demands of training. MT-NLG boasts three times the parameters of the previous largest model of its kind, showcasing how increased scale directly translates to a richer and more nuanced understanding of human language.

MT-NLG sets a new state-of-the-art standard by demonstrating unmatched accuracy in zero-, one-, and few-shot learning settings across a broad set of natural language tasks. It excels specifically in completion prediction, reading comprehension, commonsense reasoning, natural language inference, and word sense disambiguation, paving the way for advanced applications like automatic dialogue generation, translation, and code autocompletion.

Read More at the original source →