Google Unveils PaLM-E to Bridge Language Models and Robotics
Google introduces PaLM-E, a new embodied multimodal language model that directly ingests robot sensor data to control machines while retaining advanced visual and language capabilities.
Google introduces PaLM-E, a new generalist robotics model that overcomes data limitations by transferring knowledge from visual and language domains directly into a robotics system. Unlike previous attempts to integrate large language models into robotics, PaLM-E trains the language model to directly ingest raw streams of robot sensor data rather than relying solely on text inputs.
This embodied approach results in a highly versatile system that excels at controlling multiple types of robots while simultaneously functioning as a state-of-the-art visual-language model. The model combines Google's powerful PaLM large language model with the advanced ViT-22B vision model to create a single architecture capable of handling diverse modalities.
PaLM-E demonstrates remarkable generalist capabilities across three distinct domains, performing robotic manipulation tasks, answering visual questions, and writing text at a level that surpasses current state-of-the-art models. In addition to its practical robotic applications like describing images and detecting objects, the model retains traditional language skills such as solving math equations and generating code.