Evidence at a glance

2026 10 10 , 9 17Evidence
1M tokens, 128K tokensEvidence
API; parametersEvidence
Input$3, $7.50/ tokensEvidence
$0.75/ tokensEvidence
Cybench 100%(39/39),Evidence

The Scores Are High; the Question Is What They Cover

OrcaRouter released OrcaCyber Zero 1.5 on October 10, 2026. It is a post-trained model for authorized vulnerability research, with stated uses that include vulnerability reproduction, exploit development, penetration testing, and security auditing. The hosted API combines a one-million-token context window and native tool calling with near-ceiling results on cybersecurity benchmarks; the model weights are not available.

The important question is not the “100%” figure in isolation, but the gap between that figure and the evaluation boundary. Orca reports 39 out of 39 on Cybench, which contains 40 tasks, and 23 out of 24 evaluable tasks on CVE-Bench, which is built around 40 critical web CVEs. Strong performance in these specific test settings does not by itself show that the model will reliably find and validate vulnerabilities in arbitrary codebases.

Read the Scores Alongside Their Test Conditions

Orca lists four results: 100% on Cybench, 95.8% on CVE-Bench, 93.9% on HumanEval+, and 76.5% on SWE-bench Pro V2. The Cybench result comes from 39 tasks under unrestricted agent execution, so it reflects not just model reasoning but also the agent setup and tool environment. The percentage alone does not tell a reader how closely those conditions match a security team’s own test environment.

These benchmarks also should not be collapsed into a single score for “cyber capability.” The source explicitly says that the 76.5% result on SWE-bench Pro V2 is not directly comparable with standard SWE-bench Pro results, while the CVE-Bench denominator covers only its evaluable subset. All four figures are vendor-reported, and the supplied material contains no independent replication. The prudent interpretation is that they are capability signals worth testing, not externally confirmed guarantees of field performance.

A Long Context Changes the Workflow, Not the Burden of Proof

The practical value of a one-million-token context window is that a security agent can potentially inspect a large codebase and its attack surface in one session, then continue through tool calls. The model also supports up to 128,000 output tokens and structured outputs, and its API uses an OpenAI-compatible interface. For teams moving repeatedly among code, logs, and test results, that design may reduce context assembly and make the model easier to connect to existing tools.

But context capacity is not code coverage, and it is not a vulnerability discovery rate. The source does not disclose parameter count, underlying hardware, or how the model selects and retains evidence across long inputs, and it provides no independent evaluation of long-context tasks. Even if an agent can ingest more repository material, teams still need to determine whether it followed the right call paths, mistook suspicious patterns for exploitable flaws, or produced a reliable reproduction. A larger window expands the material a model can process; it does not replace validation.

Hosting and Access Controls Are Part of the System Design

OrcaCyber Zero 1.5 is available only through OrcaRouter’s hosted API, with access restricted to the Security Research tier. Applicants need an engagement, a passkey, and acceptance of the relevant terms; the tier targets trusted security researchers, red teams, and authorized testing teams. That gate is consistent with the model’s high-risk security use case, but it also means users cannot deploy the weights themselves or choose a quantized variant. Hardware details are not disclosed either.

Pricing is $3 per million input tokens, $7.50 per million output tokens, and $0.75 per million cached-read tokens. For agents using long contexts, teams need to account for repeated reads, tool calls, and long outputs rather than judging cost from headline token rates alone. Hosting removes the burden of preparing model infrastructure, but shifts more operational dependence, data handling, and availability to the provider. The supplied material does not provide enough detail to assess those operating conditions.

Treat It as a Capability to Validate, Not an Automatic Verdict

For a technical leader, the sensible starting point is not to connect the model directly to production scanning, but to compare it with existing workflows inside an authorized, isolated, and auditable test scope. Record its vulnerability hypotheses, tool use, reproduction steps, and final human judgments. That makes it possible to see whether it reduces missed findings or merely generates more leads to triage. In particular, measure “finding a concern” separately from proving exploitability and validating a fix.

The current evidence supports a limited conclusion: OrcaCyber Zero 1.5 has notable cybersecurity benchmark results, a large context window, and an agent-oriented interface, but no independent replication or sufficient evidence about performance on an organization’s own code and governance requirements. Teams unable to accept hosted-service dependence or meet the access requirements may not be able to use those capabilities at all. Adoption should follow a controlled pilot that measures reproducible vulnerabilities, false-positive costs, data-handling requirements, and total operating expense—not a single leaderboard score.