Google and Berlin Researchers Unveil PaLM-E AI Model for Autonomous Robots
Google and the Technical University of Berlin introduce PaLM-E, a massive 562-billion-parameter AI model that combines vision and language to control autonomous robots. The system understands voice commands and adapts to environmental changes in real time without needing human retraining.
Researchers from Google LLC and the Technical University of Berlin unveil PaLM-E, a massive artificial intelligence model with over 562 billion parameters designed to control autonomous robots. This multimodal system integrates AI-powered vision and language capabilities, allowing robots to understand human voice commands and execute tasks immediately without the need for constant retraining.
The model operates by viewing its immediate surroundings through the robot's camera without requiring preprocessed scene representations or human-annotated visual data. For example, when commanded to retrieve a specific item like rice chips from a drawer, PaLM-E rapidly formulates a plan based on its visual input and directs the robotic arm to complete the task fully autonomously.
PaLM-E also demonstrates an impressive ability to adapt to unexpected changes in its environment during task execution. If an item is moved by someone else, the robot sees the change, locates the object, and adjusts its plan accordingly. Additionally, the model handles complex, multi-step sequences, such as finding a sponge and bringing it to a user to clean up a spilled drink, all without requiring human guidance.