Evidence at a glance

100 Binary ExploitationEvidence
See the article text for the exact figure.Evidence
GLM-5.3 4%Evidence
Claude Mythos Preview 6%Evidence
GLM-5.3 Claude Mythos PreviewEvidence
Claude Opus 4.6 0%Evidence

This Is Not an Ordinary Model Ranking Change

Anthropic Frontier Red Team reported these findings in research on GLM-5.3 and the spread of advanced cyber capabilities. The evaluation randomly sampled 100 tasks from an internal Binary Exploitation benchmark and tested whether GLM-5.3 and Claude Mythos Preview could complete full control-flow hijacks. Simon Willison published the quotation on September 29, 2026. The material available here is a quotation of Anthropic’s findings rather than a complete evaluation report.

At first glance, this looks like a comparison between newer models. GLM-5.3 completed the target in 4% of 100 trials, while Claude Mythos Preview did so in 6%, putting GLM-5.3 below the latter on this measurement. The quotation also supplies an important historical control: Claude Opus 4.6 and GLM-5.2 reportedly had no successes under the same framing. That changes the question from which model won by a few percentage points to whether this capability has moved from unobserved to observable.

The Most Important Evidence Is the Difference Between Zero and Nonzero

The evidence can be reduced to a simple, checkable block: 100 tasks in total, a 4% success rate for GLM-5.3, 6% for Claude Mythos Preview, and 0% for both Claude Opus 4.6 and GLM-5.2. With a sample of 100, the percentages correspond roughly to four and six successful trials. The two-task gap is too small to establish a stable ranking, and neither percentage should be interpreted as a real-world probability of attack success.

Zero and nonzero results nevertheless carry different engineering implications. A zero result may mean that a model has not formed a transferable chain of techniques, or it may simply mean that this sample did not offer favorable conditions. A nonzero result shows that, for at least some combination of target, prompt, and execution conditions, the model can reach the reported full control-flow hijack metric. For an approval process, the question can no longer be limited to whether average capability improved. It must also ask whether a high-impact capability has appeared for the first time.

What a Full Control-Flow Hijack Does and Does Not Show

A full control-flow hijack is more specific than answering vulnerability questions or generating code. It concerns whether a model completed a task that gave it substantive control over a program’s execution path, rather than merely describing a class of flaw. Because the metric is closer to an operational exploitation step, even a small number of successes belongs in a model risk record more readily than a generic cybersecurity knowledge score.

The available material does not disclose the task difficulty distribution, target environments, privilege assumptions, tool configuration, or whether human intervention was involved in any trial. It also does not show whether the successes depended on particular vulnerability structures, whether they can be reproduced on another task sample, or whether a successful hijack can lead to persistence, lateral movement, or data theft. Full control-flow hijacking is therefore a concrete capability signal, not a complete attack chain from initial access to real-world compromise.

Evaluation Must Follow Deployment Permissions

For model platforms and security teams, the direct change is that the evaluation target cannot be the model in isolation. It must include the tools the model can call and the environment it can reach. A model that shows only a small number of successes in an isolated benchmark may have a shorter path from demonstrated capability to automation when deployed with repeated access to terminals, debuggers, compilers, or real code repositories. Execution privileges and opportunities for multistep trial and error may matter more to operational risk than an aggregate score outside that environment.

Pre-deployment testing should therefore reproduce the model, tools, targets, and permissions as a single system. For systems that expose sensitive code or execution environments, full control-flow hijacking can serve as a tiered gate, with the resulting permission scope and follow-on actions recorded separately. The gate does not have to mean an automatic ban. It should at least trigger stronger isolation, human approval, or restrictions on tool use instead of allowing the model to inherit the previous default assumption that zero observed successes will continue.

For Now, 4% and 6% Are Warning Signals, Not Verdicts

These figures invite two opposite mistakes. One is to treat 6% as a robust lead for Claude Mythos Preview over GLM-5.3. The other is to dismiss the capability because the success rates are low. The first mistake ignores that the sample contains only 100 randomly selected tasks. The second ignores that two earlier models had no successes while the newer models produced some. The more defensible reading is that the ranking is not stable, but the capability boundary has produced a meaningful warning signal.

The next step should not be to convert 4% or 6% directly into an organization’s real-world intrusion probability. Security teams need to establish how difficulty is distributed across the 100 tasks, whether successes transfer across environments, which tools and permissions the model requires, and how far a control-flow hijack is from later attack stages. Until those conditions are disclosed, the actionable judgment is that binary exploitation belongs in pre-deployment red-team gates, and zero observed success can no longer be treated as a safety assumption that will automatically persist.