AI Guardrails Block Legitimate Cybersecurity Researchers From Doing Their Jobs

Major AI companies like Anthropic and OpenAI are facing growing criticism from offensive cybersecurity researchers who say strict guardrails designed to prevent malicious use are also blocking legitimate security work. These researchers, whose job involves probing systems for vulnerabilities before criminals exploit them, argue that arbitrary safety restrictions hinder their ability to protect networks and users.

The tension has escalated alongside high-profile regulatory actions. In June, the U.S. government imposed export control restrictions on Anthropic's advanced models, Mythos and Fable, following reports that guardrails could be bypassed. While those restrictions have since been partially lifted, both Anthropic and OpenAI maintain vetted access programs that require researchers to apply for permission to use models with fewer cybersecurity limitations.

Security experts like veteran researcher Mark Dowd question whether large AI companies should act as gatekeepers deciding what is safe in cybersecurity. Offensive researchers emphasize that finding and exploiting vulnerabilities is essential defensive work, and overly cautious guardrails risk pushing the advantage toward adversaries who face no such limitations when using AI tools.

Read More at the original source →