DeepMind Builds Single AI Agent That Plays Games, Chats, and Controls Robots

Researchers introduce Gato, a single neural network that performs over 600 different tasks using the exact same weights. The generalist agent seamlessly switches between playing Atari games, captioning images, chatting with users, and controlling a physical robot arm.

Researchers at DeepMind introduce Gato, a multi-modal, multi-task, and multi-embodiment generalist agent built upon the principles of large-scale language modeling. Unlike traditional artificial intelligence systems designed for a single specific purpose, this single network uses the exact same set of weights to accomplish a massive variety of distinct goals.

Gato exhibits highly adaptable capabilities by evaluating its current context to determine the correct type of output. The system easily transitions from playing classic Atari games and captioning images to engaging in text-based conversations and stacking physical blocks using a real robotic arm.

The published report details the underlying architecture and the diverse dataset used to train this versatile model. By demonstrating that a single algorithm successfully outputs text, joint torques, and button presses depending on the situation, the researchers highlight a significant step toward creating truly universal artificial intelligence agents.

Read More at the original source →