Anthropic Limits Claude Mythos Release Over AI Cyberattack Fears

Anthropic restricts its new Claude Mythos Preview model due to its advanced ability to autonomously find and exploit software vulnerabilities, sparking industry-wide concern over an impending shift in cybersecurity.

Anthropic causes a major stir in the cybersecurity world by withholding its new Claude Mythos Preview model from the public due to its advanced cyberattack capabilities. Through a new initiative called Project Glasswing, Anthropic uses the model to find and patch vulnerabilities in various software systems before malicious actors can exploit them. Security expert Bruce Schneier notes that this strategic move serves as a highly effective public relations play that successfully captures media attention.

The new model demonstrates a significant leap in automated cyberattack capabilities by writing effective exploits without human involvement. Unlike older systems, Claude Mythos Preview chains together complex memory corruption bugs and operates effectively with simple one-shot prompting rather than requiring complex agent configurations. However, independent security firm Aisle replicates these same vulnerability discoveries using older, cheaper, and publicly available AI models, indicating that the underlying issue extends far beyond Anthropic's latest release.

Despite the current hype, defenders currently hold a slight advantage because AI models find vulnerabilities for patching more easily than they find and exploit them for attacks. This defensive edge is likely to shrink as increasingly powerful models become accessible to the general public. Experts agree that a fundamental sea change in cybersecurity is inevitable and approaching faster than the industry is prepared to handle.

Read More at the original source →