DeepMind Builds Gato, a Single AI Agent That Masters Diverse Tasks
DeepMind introduces Gato, a single generalist neural network that performs well across hundreds of tasks including text chats, image captioning, and robotic manipulation.
Researchers at DeepMind introduce Gato, a single generalist agent that draws inspiration from large-scale language models to handle tasks far beyond text generation. This multi-modal, multi-task, and multi-embodiment system uses a single neural network with a fixed set of weights to navigate an impressive variety of environments and challenges.
Gato seamlessly shifts between entirely different domains by treating all actions and outputs as a sequence of tokens. The same model plays Atari games, captions images, engages in conversation, and controls a physical robot arm to stack blocks. It determines the appropriate output format—whether that means generating text, issuing joint torques, or simulating button presses—based entirely on the context it receives.
The new report details the underlying architecture, the diverse training data, and the current capabilities of this versatile AI system. By demonstrating that a single policy achieves competent performance across so many distinct embodied and virtual tasks, Gato points toward a future where artificial intelligence relies on massive, general-purpose models rather than highly specialized algorithms.