Google Unveils PaLM-E to Bridge Language Models and Robotics

Google introduces PaLM-E, a new embodied multimodal language model that directly ingests robot sensor data to control machines while retaining advanced visual and language capabilities.

Google introduces PaLM-E, a new generalist robotics model that overcomes data limitations by transferring knowledge from visual and language domains directly into a robotics system. Unlike previous attempts to integrate large language models into robotics, PaLM-E trains the language model to directly ingest raw streams of robot sensor data rather than relying solely on text inputs.

This embodied approach results in a highly versatile system that excels at controlling multiple types of robots while simultaneously functioning as a state-of-the-art visual-language model. The model combines Google's powerful PaLM large language model with the advanced ViT-22B vision model to create a single architecture capable of handling diverse modalities.

PaLM-E demonstrates remarkable generalist capabilities across three distinct domains, performing robotic manipulation tasks, answering visual questions, and writing text at a level that surpasses current state-of-the-art models. In addition to its practical robotic applications like describing images and detecting objects, the model retains traditional language skills such as solving math equations and generating code.

Read More at the original source →