DeepMind Unveils Gato, a Versatile AI Agent for Over 600 Tasks
DeepMind introduces Gato, a generalist AI agent capable of performing over 600 diverse tasks including playing games, captioning images, and controlling robotic arms. Despite its broad capabilities, the system operates with significantly fewer parameters than single-task models like GPT-3.
DeepMind introduces Gato, a generalist AI agent that performs over 600 different tasks without specializing in just one area. The multi-modal, multi-task system plays video games, captions images, and controls real-world robotic arms, marking a significant step toward versatile artificial intelligence.
The system learns by example, ingesting billions of words, images, button presses, and joint torques as tokens to understand and execute various commands. Built on a Transformer architecture similar to OpenAI's GPT-3, Gato processes data from simulated and real-world environments alongside natural language and image datasets.
Remarkably, Gato achieves these diverse capabilities with only 1.2 billion parameters, which is orders of magnitude lower than single-task systems like GPT-3. Like other large language models, Gato still requires strong filters to mitigate issues such as bias and harsh language in its outputs.