Tech Giants Race to Build Ever Larger AI Models Despite Rising Costs

Following the groundbreaking release of GPT-3, artificial intelligence labs spend the year building increasingly massive neural networks. This push for sheer scale yields impressive results but raises serious questions about sustainability and bias.

The artificial intelligence industry experiences a massive shift toward building supersized models, a trend that starts with OpenAI's GPT-3. This enormous neural network demonstrates that simply increasing the number of parameters leads to striking improvements in language tasks, allowing the system to generalize across problems it never specifically trains on.

Major tech companies and top research labs eagerly embrace this "bigger is better" philosophy throughout the year, releasing their own giant models that often surpass GPT-3 in both size and capability. Firms like Microsoft and Nvidia collaborate to build massive networks, proving that scaling up existing architectures like the transformer consistently produces better performance without requiring entirely new algorithms.

However, this relentless push for scale brings significant drawbacks that the industry struggles to ignore. Training these monstrous models requires an unsustainably enormous amount of computing power, while the systems continue to absorb and mimic the toxicity and bias found in their online training data, leaving researchers to question how large these models can realistically get and at what ultimate cost.

Read More at the original source →