Microsoft and Nvidia Break Records With 530 Billion Parameter AI Model
Microsoft and Nvidia team up to build MT-NLG, a massive 530 billion parameter language model that surpasses GPT-3 in scale and accuracy across various natural language processing tasks.
Microsoft and Nvidia unveil a massive new natural language model called Megatron-Turing Natural Language Generation (MT-NLG), which boasts 530 billion parameters. This joint project breaks previous records for accuracy, reading comprehension, and reasoning, establishing itself as the largest monolithic transformer language model trained to date.
The developers train MT-NLG using a massive dataset of 270 billion tokens, combining filtered data from the open-source Pile collection with online text from Common Crawl. To handle this immense computational workload, the companies utilize a massive network of 560 Nvidia DGX A100 servers, each packed with eight high-capacity GPUs.
With three times as many parameters as the 175 billion parameter GPT-3 model, MT-NLG demonstrates unmatched performance in a wide variety of natural language tasks. The advanced AI even shows an impressive ability to infer questions beyond initial statements and accurately decipher messy or distorted text.