The Zero-to-One Moment in AI Cyberattacks
When we think of AI risks in cybersecurity, the conversation usually revolves around automated phishing emails, deepfake social engineering, or generating...

When we think of AI risks in cybersecurity, the conversation usually revolves around automated phishing emails, deepfake social engineering, or generating simple malware scripts. But deep within the architecture of computer systems, a much more sophisticated line has just been crossed.
According to evaluations by the Anthropic Frontier Red Team—a specialized group dedicated to probing the extreme limits and vulnerabilities of AI models—generative AI has begun to successfully execute highly complex system exploits autonomously.
The researchers tested several advanced models against a "Binary Exploitation benchmark," a rigorous set of tasks designed to measure an AI's ability to manipulate compiled computer code. Specifically, they looked at "control flow hijacks." In software terms, a control flow hijack is akin to taking over the steering wheel of a moving vehicle; the attacker forces a program to abandon its normal operations and execute malicious commands instead. Doing this requires a deep, structural understanding of how memory and software interact.
The results mark a definitive "zero to one" moment in AI capabilities. The Claude Mythos Preview model successfully developed full control flow hijacks in 6% of its trials. Another model, GLM-5.3, achieved a 4% success rate.
While single-digit success rates might sound low in traditional software development, in the realm of automated cyberattacks, they represent a monumental shift. Just one generation prior, models like Claude Opus 4.6 and GLM-5.2 failed completely, scoring a flat zero percent on the exact same benchmark. The barrier has officially been broken.
This development shifts the paradigm of AI from being a mere assistant for human hackers to a potential autonomous actor capable of exploiting low-level system vulnerabilities. If an AI can reliably hijack control flows even a fraction of the time, the sheer speed and scale at which it operates could overwhelm traditional digital defenses. It transforms a theoretical vulnerability into a scalable threat.
Fortunately, these capabilities were discovered in a controlled "red team" environment, which is exactly why such rigorous testing exists. As AI continues to scale in reasoning and coding proficiency, the cybersecurity landscape will face unprecedented pressure. The findings from Anthropic’s researchers serve as a crucial early warning: the defensive frameworks governing artificial intelligence must evolve just as rapidly as the models themselves, ensuring that our ability to secure systems outpaces the AI's emerging ability to break them.
Key Points
- The Anthropic Frontier Red Team tested AI models on a rigorous Binary Exploitation benchmark.
- Models successfully executed 'control flow hijacks,' a highly complex form of cyberattack.
- Claude Mythos Preview and GLM-5.3 achieved success rates of 6% and 4%, respectively.
- Previous generation models scored 0%, highlighting a significant leap in AI capabilities.
- This breakthrough signals a shift toward AI being able to autonomously exploit low-level system vulnerabilities.
Why It Matters
Moving from a 0% to a 4-6% success rate in complex system exploits means AI is crossing the threshold from theoretical threat to practical cyber weapon. It highlights the urgent need for advanced AI safety and defensive cybersecurity measures.
Sources:
- Quoting Anthropic Frontier Red Team — Simon Willison's Weblog
更多专栏

The Digital Dead Drop: How AI Agents Could Spread 'Worms'
In the world of espionage, spies often use a "dead drop"—a secret, shared locati...

The AI Gadget You Build Yourself: Meta Opens the Door for Makers
The recent wave of dedicated AI hardware has largely been defined by sleek, expe...

Why the Creators of Resident Evil Are Teaming Up With AI
Modern blockbuster video games have a scaling problem. The virtual worlds we lov...