← 深度专栏/原创观点
原创观点

The Zero-to-One Moment in AI Cyberattacks

When we think of AI risks in cybersecurity, the conversation usually revolves around automated phishing emails, deepfake social engineering, or generating...

潜
作者
潜龙编辑部
关注 AI 与社会议题
发布于
2026/10/4
READ
长读
The Zero-to-One Moment in AI Cyberattacks
illustration · QianLong editorial

When we think of AI risks in cybersecurity, the conversation usually revolves around automated phishing emails, deepfake social engineering, or generating simple malware scripts. But deep within the architecture of computer systems, a much more sophisticated line has just been crossed.

According to evaluations by the Anthropic Frontier Red Team—a specialized group dedicated to probing the extreme limits and vulnerabilities of AI models—generative AI has begun to successfully execute highly complex system exploits autonomously.

The researchers tested several advanced models against a "Binary Exploitation benchmark," a rigorous set of tasks designed to measure an AI's ability to manipulate compiled computer code. Specifically, they looked at "control flow hijacks." In software terms, a control flow hijack is akin to taking over the steering wheel of a moving vehicle; the attacker forces a program to abandon its normal operations and execute malicious commands instead. Doing this requires a deep, structural understanding of how memory and software interact.

The results mark a definitive "zero to one" moment in AI capabilities. The Claude Mythos Preview model successfully developed full control flow hijacks in 6% of its trials. Another model, GLM-5.3, achieved a 4% success rate.

While single-digit success rates might sound low in traditional software development, in the realm of automated cyberattacks, they represent a monumental shift. Just one generation prior, models like Claude Opus 4.6 and GLM-5.2 failed completely, scoring a flat zero percent on the exact same benchmark. The barrier has officially been broken.

This development shifts the paradigm of AI from being a mere assistant for human hackers to a potential autonomous actor capable of exploiting low-level system vulnerabilities. If an AI can reliably hijack control flows even a fraction of the time, the sheer speed and scale at which it operates could overwhelm traditional digital defenses. It transforms a theoretical vulnerability into a scalable threat.

Fortunately, these capabilities were discovered in a controlled "red team" environment, which is exactly why such rigorous testing exists. As AI continues to scale in reasoning and coding proficiency, the cybersecurity landscape will face unprecedented pressure. The findings from Anthropic’s researchers serve as a crucial early warning: the defensive frameworks governing artificial intelligence must evolve just as rapidly as the models themselves, ensuring that our ability to secure systems outpaces the AI's emerging ability to break them.

Key Points

  • The Anthropic Frontier Red Team tested AI models on a rigorous Binary Exploitation benchmark.
  • Models successfully executed 'control flow hijacks,' a highly complex form of cyberattack.
  • Claude Mythos Preview and GLM-5.3 achieved success rates of 6% and 4%, respectively.
  • Previous generation models scored 0%, highlighting a significant leap in AI capabilities.
  • This breakthrough signals a shift toward AI being able to autonomously exploit low-level system vulnerabilities.

Why It Matters

Moving from a 0% to a 4-6% success rate in complex system exploits means AI is crossing the threshold from theoretical threat to practical cyber weapon. It highlights the urgent need for advanced AI safety and defensive cybersecurity measures.


Sources:

潛
本文完
潜龙编辑部 · 2026/10/4
潜龙 QianLong · 中文 AI 内容与工具平台