DeepMind Builds Single AI Agent That Plays Games, Chats, and Controls Robots
DeepMind introduces Gato, a multi-modal AI agent that uses a single neural network to perform hundreds of diverse tasks ranging from playing Atari games to controlling a real robotic arm.
DeepMind researchers introduce Gato, a groundbreaking generalist agent that applies large-scale language modeling techniques to physical and digital environments. Unlike traditional AI systems that specialize in a single task, this single network uses the exact same weights to execute an impressive variety of actions based on the context it receives.
The system operates as a multi-modal, multi-task, and multi-embodiment policy that seamlessly switches between different output types. Depending on the situation, Gato decides whether to output text for chatting, joint torques for controlling a real robot arm, button presses for playing Atari games, or image captions.
By training on a massive and diverse dataset, the model demonstrates that a single architecture achieves competent performance across hundreds of different tasks without requiring separate, highly specialized systems. This report documents the current capabilities of Gato and highlights a significant step toward creating truly versatile artificial intelligence.