AI Industry Embraces Supersized Models Despite Massive Costs

The tech industry spends 2021 building increasingly massive AI models, driven by the belief that sheer size leads to better performance. This trend raises serious questions about sustainability, bias, and the ultimate limits of scaling.

The artificial intelligence industry experiences a massive shift in 2021 as tech giants and top research labs race to build increasingly enormous models. This trend originates from OpenAI's GPT-3, a groundbreaking neural network that achieves uncanny language abilities not through new algorithms, but through sheer scale. By packing more parameters into the system, developers discover that bigger models generalize better across tasks they are not explicitly trained to perform.

Multiple companies, including Microsoft and Nvidia, release their own supersized models throughout the year, with several surpassing GPT-3 in both size and capability. Researchers openly admit they originally thought they needed entirely new ideas to advance the field, but they instead find that simply adding more computing power and parameters drives significant performance gains. Major tech players continue to hype this approach, claiming there is seemingly no end in sight to the benefits of hyperscaling.

However, this relentless pursuit of size brings significant downsides that the industry struggles to address. Training these monster models requires an unsustainably enormous amount of computing power, raising serious environmental and financial concerns. Furthermore, because these systems learn from vast troves of online text, they inevitably absorb and replicate the toxicity and bias present in that data, leaving developers to grapple with the ethical consequences of their massive creations.

Read More at the original source →