Tech Giants Race to Build Ever Larger AI Models Despite Rising Costs

The AI industry experiences a massive shift toward building supersized neural networks, a trend sparked by OpenAI's GPT-3. Researchers continue to push the limits of scale, raising questions about the future costs and capabilities of these enormous models.

The artificial intelligence industry experiences a massive shift toward building supersized neural networks, a trend sparked by OpenAI's GPT-3 in 2020. This year sees a proliferation of even larger models from multiple tech firms and top AI labs that surpass GPT-3 in both size and ability. These massive networks achieve striking jumps in performance and generalization not through new algorithms, but through sheer scale.

Researchers measure the size of these trained neural networks by their parameter count, which represents the values tweaked during the training process. Tech giants like Microsoft and Nvidia collaborate to build enormous models like the Megatron-Turing NLG. Experts note that hyperscaling these models leads to better performance with seemingly no end in sight.

However, this push toward bigger models comes with significant drawbacks. Training these networks requires an unsustainably enormous amount of computing power, raising serious questions about the financial and environmental costs. Additionally, these massive models continue to mimic the bias and toxicity found in the vast amounts of online text used to train them.

Read More at the original source →