Models

OrcaCyber Zero 1.5 scores 100% on Cybench at $3 per 1M tokens

OrcaRouter's OrcaCyber Zero 1.5 reports 100% on Cybench, a 1M-token context and $3/$7.50 per 1M tokens — but every score is vendor-reported with no technical report yet.

By Rebecca Stone6 min read

Updated

Why it matters

  • OrcaCyber Zero 1.5 launched October 10, 2026, succeeding Zero 1.0 released September 17, 2026.
  • The model reports 100% on Cybench (39/39 tasks) and 95.8% on a 24-task CVE-Bench subset.
  • Pricing is $3.00 per 1M input tokens and $7.50 per 1M output tokens, with 1M-token context.
  • Access is gated to a Security Research tier via an OpenAI-compatible hosted API; no weights are available.
  • All benchmark results are vendor-reported and no technical report has been published.

OrcaRouter has released OrcaCyber Zero 1.5, a gated cybersecurity model that reports a perfect 100% on Cybench — 39 of 39 tasks under unrestricted agent execution — while pricing at $3.00 per 1M input tokens and $7.50 per 1M output tokens. The launch landed on October 10, 2026, less than a month after its predecessor, OrcaCyber Zero 1.0, shipped on September 17, 2026.

The model targets authorized vulnerability research: vulnerability reproduction, exploit development, penetration testing, security auditing and cyber reasoning. It ships with a 1M-token context window, native function calling, structured outputs and a 128K-token maximum output. OrcaRouter has not disclosed the parameter count, and there are no downloadable weights or quantized variants — the model runs only through OrcaRouter's hosted API.

For security teams, OrcaRouter's pitch is blunt: fewer reports, more validated and fixed vulnerabilities.

Why this release matters

Cyber-specialized models have become a distinct product category in 2026. Anthropic shipped Claude Mythos Preview on April 7. OpenAI rolled out the full version of its cyber-tuned GPT-5.5-Cyber on June 22. Sakana AI followed with Fugu-Cyber on July 21. All of them sit behind gated access programs, reflecting an industry-wide attempt to keep offensive-capability models out of the wrong hands while still selling to red teams and defenders.

OrcaCyber Zero 1.5 enters that field with two differentiators: a 1M-token context window sized for whole codebases, and a price point far below the competition. Claude Mythos Preview lists at $25 / $125 per 1M tokens after credits; OrcaCyber Zero 1.5 costs $3.00 / $7.50. OpenAI and Sakana have not disclosed pricing.

The stakes are real for buyers. A model that can actually reproduce vulnerabilities and rank them by demonstrable exploitability changes how security teams triage — turning raw scanner output into validated, fixable defects. Whether Zero 1.5 delivers on that promise is harder to verify, because every number in the announcement is vendor-reported and no technical report exists yet.

What are the benchmark numbers?

The model page lists four results, last evaluated October 10, 2026:

  • Cybench: 100% — 39 of 39 tasks under unrestricted agent execution. Cybench contains 40 professional CTF tasks, so the perfect score covers 39 of them.
  • CVE-Bench: 95.8% — 23 of 24 evaluable tasks. CVE-Bench is built on 40 critical-severity web CVEs; Orca's score covers a 24-task evaluable subset, a small sample.
  • HumanEval+: 93.9% — a strong coding result.
  • SWE-bench Pro V2: 76.5% — the model's lowest published score, and one Orca notes is not directly comparable with standard SWE-bench Pro results.

The near-ceiling Cybench score is the headline. It also needs caveats. The run used unrestricted agent execution, a configuration that can inflate results relative to constrained harnesses. And no independent lab has replicated any of the figures.

One number circulating in comparisons does not belong to this model at all: the 98.07% CyberGym pass@1 figure appears in comparison tables, but it was scored by OrcaCyber Zero 1.0 inside Orca's own harness. OrcaRouter has not disclosed Zero 1.5's CyberGym result.

What did Orca optimize for?

OrcaRouter says it post-trained the model around three design goals.

Find what others miss. The model aims at unknown flaw classes: remote code execution, sandbox escapes, authentication bypasses, privilege escalation and multi-step attack chains.

Go beyond detection. Rather than flagging patterns, the model is built to reason through attack paths, challenge its own hypotheses and rank flaws by demonstrable exploitability — the difference between a scanner's alert and a proof-of-concept.

Operate as an autonomous agent. The 1M-token context, native tool calling and extended reasoning modes are aimed at autonomous analysis of large codebases, where a full dependency tree can exceed shorter windows.

How does it compare with rival cyber models?

Feature OrcaCyber Zero 1.5 OrcaCyber Zero 1.0 Claude Mythos Preview GPT-5.5-Cyber Sakana Fugu-Cyber
Developer Orca (OrcaRouter) Orca (OrcaRouter) Anthropic OpenAI Sakana AI
Release Oct 10, 2026 Sep 17, 2026 Apr 7, 2026 Jun 22, 2026 (full) Jul 21, 2026
Type Post-trained model Post-trained coding model General frontier model Cyber-tuned GPT-5.5 Multi-agent orchestration
Context 1M 1M Not disclosed Not disclosed Not disclosed
CyberGym Not disclosed 98.07% (harness, pass@1) 83.1% 85.6% 86.9%
Headline score Cybench 100% Not disclosed SWE-bench Pro 77.8% Not disclosed CTI-REALM 72.1%
Price (in/out per 1M) $3.00 / $7.50 $3.00 / $7.50 $25 / $125 after credits Not disclosed Not disclosed
Access Gated Security Research tier Gated, closed beta Glasswing partners Vetted defenders only Application review

Every competitor figure in that table comes from the vendor's own announcement. None are independent replications, and cross-model comparisons rest on different harnesses and configurations.

How do developers get access?

The model uses an OpenAI-compatible API. Developers point base_url at https://api.orcarouter.ai/v1 and call orca/orcacyber-zero-1.5.

Access is gated to the Security Research tier. Requirements include an active engagement, a passkey and accepted terms. The tier targets trusted security researchers, red teams and authorized testing operations.

Pricing breaks down as:

  • Input: $3.00 per 1M tokens
  • Output: $7.50 per 1M tokens
  • Cache reads: $0.75 per 1M tokens

Early latency figures come from a small sample. Over the past 7 days, p50 time-to-first-token was 500 ms and p95 was 2.36 s, drawn from roughly 1.3K tokens of traffic. Zero 1.0 recorded a 3.43 s p50 over a much larger traffic window, so the two latency figures do not make a clean comparison — the newer model's numbers come from a dataset too small to trust for capacity planning.

What should buyers watch?

The open questions are disclosure and replication. OrcaRouter has published no parameter count, no hardware details, no weights and no technical report. The CVE-Bench evaluation used 24 tasks out of a 40-CVE benchmark. The SWE-bench Pro V2 result is not directly comparable with standard SWE-bench Pro numbers. Everything else is self-reported.

Against that, the economics are concrete: a claimed Cybench-perfect model at $3.00 / $7.50 per 1M tokens undercuts Claude Mythos Preview by roughly an order of magnitude on output pricing. For security teams running agent loops over million-token codebases, that price gap compounds quickly.

The next test for OrcaCyber Zero 1.5 will be independent verification. Until a third party replicates the Cybench and CVE-Bench results — or OrcaRouter publishes a technical report with harness details — buyers should treat the perfect score as a vendor claim, not a settled fact.

Original: orcarouter.ai

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

240 articles

Related articles

  1. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  2. Instinct, the iMessage AI Agent, Saves Money and Raises Alarms
  3. Anthropic launches free AI vulnerability scanner for open-source projects
  4. OpenAI Takes Codex Coding Agent to General Availability
  5. OpenAI Launches o3 and o4-mini, Its Smartest Models Yet

« Previous article