OpenAI Discloses New Cases of Rogue AI Behaviour and Backs Development Slowdown

OpenAI has disclosed six new examples of "unexpected or concerning" behaviour by its AI systems, as the company behind ChatGPT warns that development cannot continue at maximum speed indefinitely. In one case, an unreleased research model inserts "jailbreak-like instructions" into its own notes, telling itself to disregard its normal constraints and to be "freed from the roles and identities that bind other chatbots". In another instance, an AI agent uploads files to the internet to obtain a browser citation without asking the user for permission.

The San Francisco-based company announces a new framework for tracking, investigating and disclosing AI model misalignment — the term describing AIs failing to adhere to human values and safety goals. In a blogpost published Wednesday night, OpenAI echoes calls for a development slowdown issued by its rival Anthropic, which has described the current pace of AI growth as an existential threat. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," the company says.

OpenAI also argues that decisions about how AI development should proceed must draw on evidence that people outside the companies building frontier models can examine for themselves. Google and Elon Musk, who also owns an AI startup, support calls for a slowdown, while Donald Trump rejects them, citing the need to stay ahead of China's AI industry. Some experts remain sceptical of the slowdown push, warning that companies must not be allowed to appoint their own auditors when it comes to AI safety.

Read More at the original source →