Tech Giants Race to Build Ever Larger AI Models Despite High Costs
The AI industry experiences a massive shift toward building supersized neural networks, a trend sparked by OpenAI's GPT-3. Tech companies continue to push the limits of scale, raising questions about the future costs and sustainability of these enormous models.
The artificial intelligence industry experiences a massive shift toward building supersized neural networks, a trend that sparks directly from OpenAI's release of GPT-3. This enormous language model proves that simply increasing the size of a neural network leads to striking improvements in its ability to generalize across tasks it is not specifically trained to perform.
Throughout the year, top tech firms and AI labs release a proliferation of large models that surpass GPT-3 in both size and capability. Companies like Microsoft and Nvidia collaborate to build massive networks, as researchers note that this hyperscaling of AI models leads to better performance with seemingly no end in sight.
However, this pursuit of sheer scale comes with significant drawbacks and unanswered questions. Training these monstrous models requires an unsustainably enormous amount of computing power, and the systems frequently mimic the bias and toxicity found in their online training data, leaving the industry to grapple with how large these models will ultimately get and at what cost.