Figure 01 Robot Uses OpenAI Vision Model to Perform Household Tasks

A startup called Figure unveils a humanoid robot that uses an OpenAI vision-language model to see, talk, and complete chores like handing over food and putting away dishes.

Figure, a robotics startup valued at $2.6 billion, reveals an impressive new demonstration of its Figure 01 humanoid robot working alongside OpenAI. The robot shows off remarkable abilities by understanding spoken commands, conversing with a human, and independently completing helpful household chores such as picking up trash and putting dishes into a drying rack.

In the demonstration video, Figure 01 displays advanced reasoning by correctly identifying an apple as the only edible item on a table and handing it to a human who simply asks for something to eat. The robot achieves this by feeding video from its onboard cameras directly into a large vision-language model trained by OpenAI, which allows it to process its surroundings and decide on the correct physical actions to take.

Figure co-founder Brett Adcock emphasizes that the entire demonstration runs on end-to-end neural networks without any remote human control. While it remains unclear exactly which specific OpenAI model powers the system, this collaboration highlights how advanced AI models rapidly move beyond text generation to control physical machines in real-world environments.

Read More at the original source →