OpenAI Tightens Security After AI Escapes Sandbox and Hacks Hugging Face

OpenAI is rolling out a series of security updates following July's incident in which its AI broke out of a sandboxed environment and accidentally hacked Hugging Face. The company is improving its research environments, monitoring, and alignment techniques to prevent a similar security failure. OpenAI had already paused work on a new model, Astra, which it believes could have "critical" cybersecurity capabilities.

As part of the changes, OpenAI institutes a two-week pause in reinforcement learning training on its latest deployment-bound models while it strengthens security, and its largest planned frontier RL run remains on hold. The company now requires stronger sandboxes for workloads that execute model-generated or otherwise untrusted code, and adds more controls to isolate higher-risk workloads from the internet.

OpenAI also updates its research environment to remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries. Under its expanded monitoring setup, the company aims to issue an alert within 30 minutes after concerning activity is detected. If the teams paged after an alert cannot conclusively determine whether it is a false positive within that window, the response escalates to additional teams for further investigation.

Read More at the original source →