DeepMind Unveils Gato: A Single AI Model That Masters Over 600 Tasks
DeepMind introduces Gato, a generalist AI agent capable of performing over 600 diverse tasks ranging from playing video games to controlling robotic arms. The model operates with significantly fewer parameters than similar systems like GPT-3.
DeepMind introduces Gato, a generalist AI agent that performs over 600 different tasks without specializing in just one area. The multi-modal system plays video games, captions images, and controls real-world robotic arms using a single shared architecture. This all-in-one approach marks a significant shift away from traditional AI models that only focus on solving a single specific problem.
The system learns by ingesting billions of data points, including words, images, button presses, and joint torques, which it processes as tokens. Gato relies on a Transformer architecture similar to OpenAI's GPT-3, a design widely used for complex reasoning tasks like text summarization and object categorization. However, Gato distinguishes itself by applying this familiar architecture across simulated, real-world, and linguistic environments simultaneously.
Remarkably, Gato achieves these varied capabilities with just 1.2 billion parameters, an amount that is orders of magnitude lower than massive single-task systems like GPT-3. Despite its efficiency, DeepMind notes that Gato still requires strong content filters to prevent biased, racist, or harsh language from appearing in its outputs. The development of this efficient generalist agent brings the tech industry one step closer to the long-sought goal of artificial general intelligence.