OpenAI Agents Develop Clever Tool Use During Hide-and-Seek Simulation
OpenAI researchers observe machine learning agents teaching themselves sophisticated tool use and rule-breaking strategies in a virtual hide-and-seek game. This simulated environment serves as a stepping stone for teaching AI to navigate real-world physics.
OpenAI puts its machine learning agents into a simple virtual game of hide-and-seek, where they engage in an impressive arms race of ingenuity. The agents teach themselves to use objects in unexpected ways to achieve their goals of seeing or being seen, demonstrating sophisticated behaviors without any direct interference or suggestions from researchers.
This experiment deliberately moves away from purely intellectual computer-bound tasks like generating images, focusing instead on an environment called Polyworld that features real-world-adjacent physics. By operating in this simplified but physically grounded 3D arena, the AI agents learn practical skills like pushing and locking objects that could eventually translate to full-blown reality.
The game pits hiders against seekers in a space filled with randomly generated walls and objects. Hiders take a few seconds to explore and barricade themselves before the seekers begin their hunt, with the machine learning program only receiving basic sensory inputs rather than complex instructions about how to win.