Tech Giants Race to Build Ever Larger AI Models Despite Rising Costs

The AI industry experiences a massive shift toward building supersized neural networks, a trend sparked by OpenAI's GPT-3. Tech companies continue to push the limits of scale, raising questions about the future costs and sustainability of these enormous models.

The artificial intelligence industry experiences a massive shift toward building supersized neural networks, a trend that sparks with the release of OpenAI's GPT-3 in 2020. This groundbreaking model demonstrates that simply increasing the number of parameters leads to striking improvements in language tasks, convincing researchers that sheer scale often outperforms new algorithmic ideas.

In 2021, multiple tech firms and top AI labs embrace this bigger-is-better philosophy by releasing their own enormous models that frequently surpass GPT-3 in both size and capability. Companies like Microsoft and Nvidia collaborate to build massive networks like the Megatron-Turing NLG, showing that the industry sees seemingly no end in sight for the performance gains achieved through hyperscaling.

However, this rapid growth brings significant challenges as these monster models require an unsustainably enormous amount of computing power to train. Furthermore, they continue to mimic the bias and toxicity found in their online training data, forcing the tech world to weigh the impressive capabilities of these giant networks against their high financial and ethical costs.

Read More at the original source →