Tech Giants Race to Build Ever Larger AI Models Despite Rising Costs
The AI industry experiences a massive shift toward building supersized neural networks, a trend sparked by OpenAI's GPT-3. Researchers continue to push the limits of scale to achieve better performance, raising questions about sustainability and bias.
The artificial intelligence industry experiences a massive shift toward building supersized neural networks, a trend that sparks heavily with the release of OpenAI's GPT-3. This massive model demonstrates an uncanny grasp of human language, generating convincing text and completing code by relying on sheer size rather than new algorithms.
Throughout the year, multiple tech firms and top AI labs release their own enormous models that surpass GPT-3 in both size and capability. Researchers discover that simply adding more parameters leads to striking jumps in performance and an improved ability to generalize across tasks without specific training.
However, this relentless pursuit of scale brings significant challenges to the forefront. Training these monstrous networks requires an unsustainably enormous amount of computing power, and the models frequently mimic the bias and toxicity found in their online training data, leaving the industry to question how large these systems will ultimately become.