Safety & Security

OpenAI cyber capability scores jump from 27% to 76% in three months

OpenAI's GPT-5.1-Codex-Max now scores 76% on capture-the-flag challenges, up from 27% on GPT-5 in August 2025. The lab is preparing Preparedness Framework safeguards, an Aardvark beta, and trusted access for defenders.

Strengthening cyber resilience as AI capabilities advance
Strengthening cyber resilience as AI capabilities advanceAI-generated
By Marcus Bennett4 min read

Updated

Why it matters

  • Capture-the-flag scores rose from 27% on GPT-5 (August 2025) to 76% on GPT-5.1-Codex-Max (November 2025)
  • OpenAI defines the 'High' Preparedness Framework cyber tier as zero-day remote exploits or stealthy enterprise/industrial intrusion assistance
  • Aardvark, OpenAI's agentic security researcher, entered private beta and has identified novel CVEs in open-source codebases
  • OpenAI is designing a tiered trusted access program for qualifying cyberdefense users, with boundaries still under evaluation
  • A new Frontier Risk Council will advise on the line between responsible capability and misuse, starting in cybersecurity

OpenAI's flagship models have nearly tripled their performance on capture-the-flag cybersecurity challenges in roughly three months, climbing from 27% on GPT-5 in August 2025 to 76% on GPT-5.1-Codex-Max in November 2025. The jump, disclosed in a company post on cyber resilience, frames the rapid gains as both a defensive asset and a dual-use risk that the lab says it now plans for at its highest internal threat tier.

The trajectory matters because cybersecurity sits among the most contested dual-use domains in commercial AI. Defensive and offensive workflows share the same underlying knowledge; the same exploit that helps a defender patch a bug can help an attacker weaponize it. OpenAI says it is "planning and evaluating as though each new model could reach 'High' levels of cybersecurity capability, as measured by our Preparedness Framework."

What does "High" cyber capability mean?

OpenAI defines that tier as models that can either develop working zero-day remote exploits against well-defended systems, or meaningfully assist with complex, stealthy enterprise or industrial intrusion operations aimed at real-world effects. The company frames the assumption as a planning baseline, not a forecast.

The framing has practical consequences. Once a model crosses the High threshold, OpenAI would be expected to apply its most stringent Preparedness Framework controls, including tighter deployment gating and external review. The lab is pre-building the safety scaffolding for a capability it expects frontier models to reach on the current trajectory.

How is OpenAI layering its defenses?

OpenAI is taking a defense-in-depth approach that rejects single-point safeguards. At the foundation sit access controls, infrastructure hardening, egress controls, and monitoring. The company layers detection and response systems on top, plus dedicated threat intelligence and insider-risk programs.

On the model layer, OpenAI is training frontier systems to refuse or safely respond to harmful cyber requests while staying helpful for educational and defensive use cases. Detection systems monitor products for malicious activity. When prompts look unsafe, OpenAI may block output, route requests to safer or less capable models, or escalate for human enforcement that weighs legal requirements, severity, and repeat behavior.

The third layer is end-to-end red teaming. OpenAI says it works with expert organizations that try to bypass every layer of its defenses the way a "determined and well-resourced adversary" might. The goal: surface gaps before an attacker does.

What is Aardvark?

Aardvark is OpenAI's agentic security researcher, now in private beta. The tool scans codebases for vulnerabilities and proposes patches that maintainers can adopt quickly. According to OpenAI, Aardvark has already identified novel CVEs in open-source software by reasoning over entire codebases.

The company plans to offer free coverage to select non-commercial open-source repositories. The move responds to a long-standing concern: that AI-assisted vulnerability discovery could outpace the open-source community's ability to patch, creating a defender's dilemma. OpenAI is positioning Aardvark as a force multiplier for the maintenance side of that equation.

Who will get trusted access?

OpenAI is preparing a trusted access program that will give qualifying cyberdefense users and customers tiered access to enhanced capabilities in the lab's latest models. The boundaries — which capabilities get broad release and which require tiered restrictions — remain under design.

The program reflects a recurring tension in frontier AI deployment. A blanket restriction on cyber capabilities would blunt the defensive value of the same models. A blanket release would hand attackers leverage. Tiered access is OpenAI's attempt to thread the needle, though the company concedes the right shape of the program is still in flux.

How will outside expertise shape the rollout?

Two new structures will guide the work. The Frontier Risk Council will advise OpenAI on the boundary between useful, responsible capability and potential misuse, starting with cybersecurity before expanding into other frontier domains. Separately, OpenAI will continue working through the Frontier Model Forum — a nonprofit backed by leading AI labs — to align on threat models and best practices across the industry.

The company is also engaging with external teams to develop cybersecurity evaluations. OpenAI says it hopes "an ecosystem of independent evaluations will further help build a shared understanding of model capabilities." Independent benchmarks matter because lab-internal evaluations can miss adversarial use cases the lab itself has not imagined.

What are the stakes?

The cyber capability curve is steepening at the same moment that more of the world's critical infrastructure is being managed by software. OpenAI is making two implicit bets: that its safeguards can keep pace with the capability curve, and that the defensive community can absorb the leverage faster than attackers can.

Both bets are testable. The trusted access program, Aardvark's open-source rollout, and the Frontier Risk Council's first recommendations will all signal whether the lab's defense-in-depth approach scales to a world where frontier models can credibly discover zero-days. If the safeguards work, defenders get a generational tool. If they do not, the same tools land in attacker hands. OpenAI is building for the first outcome while planning for the second.

Source: OpenAI News

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

178 articles

Related articles

  1. OpenAI Rolls Out GPT-5.4-Cyber to Vetted Defenders
  2. OpenAI ships GPT-5.4 Thinking with first High-tier cyber mitigations
  3. OpenAI Flags GPT-5.3-Codex as High Cybersecurity Risk
  4. OpenAI ships GPT-5.5-Cyber, tiers access for defenders
  5. OpenAI says Astra hits critical cyber capability threshold

« Previous article