Safety & Security

OpenAI details how external testers probe its frontier models

OpenAI says it gave some external testers chain-of-thought access and unmitigated models, and confirmed it pays assessors — with publication review attached.

Strengthening our safety ecosystem with external testing
Strengthening our safety ecosystem with external testingElogia Marketing4eCommerce / Openverse
By Rebecca Stone4 min read

Updated

Why it matters

  • OpenAI gave some external assessors direct chain-of-thought access and models with fewer mitigations during pre-deployment testing of GPT-5
  • Third-party assessors sign NDAs and OpenAI reviews their publications for confidentiality and factual accuracy before release
  • OpenAI pays all third-party assessors via direct payment or API credits, but states no payment is contingent on assessment results

OpenAI has laid out, in unusual detail, how it lets outside organizations test its frontier models before deployment — including giving some assessors direct access to model chain-of-thought reasoning and versions with safety mitigations stripped out.

In a blog post titled "Strengthening our safety ecosystem with external testing," OpenAI described three forms of third-party collaboration it has run since the launch of GPT-4: independent evaluations of frontier capability and risk areas such as biosecurity, cybersecurity, AI self-improvement and scheming; methodology reviews of how OpenAI evaluates and interprets risk; and subject-matter expert (SME) probing, where experts run their own real-world tasks against a model and return structured survey input.

The disclosures matter because independent evaluation of frontier models has become a central governance question. Regulators in the U.S. and U.K. have built AI safety institutes around the premise that outside assessors can check labs' internal safety claims, and OpenAI is arguing that a funded, independent assessor ecosystem is now a structural requirement for frontier AI.

Chain-of-thought access and unmitigated models

For GPT-5, OpenAI said it coordinated external capability assessments across risk areas including long-horizon autonomy, scheming, deception and oversight subversion, wet lab planning feasibility, and offensive cybersecurity. These ran alongside OpenAI's own Preparedness Framework evaluations and used external benchmarks such as METR's time horizon evaluation and SecureBio's Virology Capabilities Test (VCT).

To enable the work, OpenAI provided secure access to early model checkpoints, zero-data retention where needed, and models with fewer mitigations. Cybersecurity and biosafety testers evaluated models both with and without safety mitigations to probe underlying capabilities. Several organizations received direct chain-of-thought access, which OpenAI said "allowed assessors to identify cases of sandbagging or scheming behavior that might only be discernible through reading the chain-of-thought." Access came with security controls the company says it continues to update as capabilities evolve.

Methodology review instead of replication

For the gpt-oss release, OpenAI used adversarial fine-tuning to estimate whether a malicious actor could push the open-weight model to High capability in bio or cyber areas under its Preparedness Framework. Rather than ask outsiders to repeat that resource-intensive work, OpenAI invited assessors to review its methods and results over a multi-week process. Their feedback changed the final adversarial fine-tuning process, and OpenAI documented which recommendations it adopted — and why it rejected others — in the gpt-oss system card and accompanying paper. METR published its own review of the methods and evidence, highlighting decision-relevant gaps that were addressed through the feedback loop.

OpenAI framed this as a template for cases where the infrastructure required for worst-case experiments exists only inside major AI labs, making direct external replication impractical.

Expert probing on bio risk

For ChatGPT Agent and GPT-5, OpenAI invited a subject-matter expert panel to run their own end-to-end bio scenarios against a helpful-only model. The experts scored how much the model could uplift someone at their own level of expertise versus a less experienced novice, stress-testing OpenAI's "novice uplift" claims under realistic workflows. The exercise fed into deployment decisions and appeared in both system cards. OpenAI distinguishes this from red teaming, which targets specific safeguards rather than overall capability.

Money, NDAs and publication rights

OpenAI also published contract excerpts governing its assessor relationships. Third parties sign non-disclosure agreements, and OpenAI reviews and approves their publications for confidentiality and factual accuracy before release — a gatekeeping role that will likely draw scrutiny from those who want fully independent publication. Published work to date includes METR's GPT-5 report, Apollo Research's report on OpenAI o1, and Irregular's GPT-5 assessment.

The company confirmed it pays all of its third-party assessors, through direct payment or subsidized API credits, though some organizations decline on principle. "No payment is ever contingent on the results of a third party assessment," OpenAI wrote.

Looking ahead, OpenAI said credible third-party evaluation requires specialized expertise, stable funding and methodological rigor, and that continued investment in qualified assessor organizations will be essential to keep assessments aligned with advancing model capabilities. The company noted these assessments sit alongside other external mechanisms, including its work with U.S. CAISI and UK AISI, collective alignment projects, and advisory groups such as its Global Physician Network.

Original: cdn.openai.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Publishes Frontier Governance Framework for AI Risk
  2. OpenAI Launches Initiative for Democratic AI Oversight in National Security
  3. OpenAI ships GPT-5.5-Cyber, tiers access for defenders
  4. White House Wants First Look at New OpenAI and Anthropic Models
  5. OpenAI Lays Out Vision for AGI That "Benefits Everyone"

Next article »