Models

OpenAI Releases gpt-oss-120b and gpt-oss-20b Under Apache 2.0

OpenAI ships gpt-oss-120b and gpt-oss-20b under Apache 2.0 with full chain-of-thought and tool use, after adversarial fine-tuning tests found no High-risk capabilities.

Gpt-oss-120b & gpt-oss-20b Model Card
Gpt-oss-120b & gpt-oss-20b Model CardRoberts, David M. R. / Openverse
By Elena Vasquez7 min read

Updated

Why it matters

  • gpt-oss-120b and gpt-oss-20b are open-weight, text-only reasoning models released under the Apache 2.0 license and OpenAI's gpt-oss usage policy.
  • Adversarial fine-tuning tests reviewed by OpenAI's Safety Advisory Group found gpt-oss-120b did not reach High capability in Biological and Chemical Risk or Cyber risk.
  • OpenAI says existing open models' default performance comes near to matching adversarially fine-tuned gpt-oss-120b on most biological capability evaluations, so the release does not significantly advance the open-model frontier.

OpenAI has released gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models available under the Apache 2.0 license and the company's gpt-oss usage policy. The announcement arrives via the models' official documentation, and it marks a significant shift in how one of the industry's most prominent AI developers distributes capable systems: the weights are public, the license is permissive, and the safety analysis is published alongside the code.

The move matters because open-weight models change the risk equation entirely. Once released, OpenAI states plainly in the model card, "determined attackers could fine-tune them to bypass safety refusals or directly optimize for harm without the possibility for OpenAI to implement additional mitigations or to revoke access." That is a fundamentally different proposition from API-served models, where the company can monitor usage, update guardrails, and cut off bad actors. With gpt-oss, the genie is out of the bottle by design.

What the models do

Both models are text-only. They are compatible with OpenAI's Responses API and are built for agentic workflows — the fast-growing category of AI systems that plan, call tools, and act with limited supervision. The model card highlights three capabilities: strong instruction following, tool use including web search and Python code execution, and reasoning.

One detail stands out for developers: the models let users adjust the reasoning effort for tasks that don't require complex reasoning. That flexibility matters in production, where heavy chain-of-thought computation wastes latency and compute on simple queries. Users can dial reasoning up for hard problems and down for routine ones.

The models also provide full chain-of-thought, meaning the reasoning steps are visible rather than hidden. That is a meaningful transparency choice. It allows developers to audit how the model arrived at an answer, debug failures, and customize behavior. Structured Outputs are supported as well, which lets applications force model responses into strict schemas — a practical requirement for piping model output into databases, APIs, and downstream software.

OpenAI says it developed the models with feedback from the open-source community. The models are customizable, according to the documentation, which positions them for the fine-tuning and adaptation work that enterprises and independent developers routinely perform on open-weight foundations.

Why OpenAI calls it a model card, not a system card

The document's framing carries weight. OpenAI explicitly labels the release a "model card" rather than a "system card" — the term the company uses for its flagship proprietary releases — because, in its words, "the gpt-oss models will be used as part of a wide range of systems, created and maintained by a wide range of stakeholders."

That distinction is doing real analytical work. A system card describes protections that ship with a finished product. A model card describes a component that others will assemble into products OpenAI neither controls nor observes. The models are designed to follow OpenAI's safety policies by default, the company notes, but "other stakeholders will also make and implement their own decisions about how to keep those systems safe."

For developers and enterprises, the implication is direct: in some contexts, they will need to implement extra safeguards to replicate the system-level protections built into models served through OpenAI's API and products. The default safety posture travels with the weights; the full defensive stack does not.

This is the central trade-off of the release. Open weights democratize access to frontier-adjacent reasoning capabilities and enable customization, transparency, and vendor independence. They simultaneously offload a portion of the safety burden onto whoever deploys them. The model card acknowledges both sides without editorializing.

The safety evaluations

OpenAI ran scalable capability evaluations on gpt-oss-120b and reported a clear bottom line: the default model does not reach the company's indicative thresholds for High capability in any of the three Tracked Categories of its Preparedness Framework — Biological and Chemical capability, Cyber capability, and AI Self-Improvement.

The Preparedness Framework is OpenAI's internal system for assessing catastrophic risks from frontier models. Clearing the "not High" bar in all three tracked categories is the analytical basis on which the company judged the release acceptable.

But OpenAI went further than evaluating the default model. The team investigated two questions that get at the specific dangers of open weights.

First: could adversarial actors fine-tune gpt-oss-120b to reach High capability in the Biological and Chemical or Cyber domains? To answer this, OpenAI simulated an attacker. The team adversarially fine-tuned gpt-oss-120b in both categories — deliberately training it toward dangerous capability rather than away from it. OpenAI's Safety Advisory Group (SAG) reviewed the testing. Its conclusion: even with robust fine-tuning that leveraged OpenAI's own field-leading training stack, gpt-oss-120b did not reach High capability in Biological and Chemical Risk or Cyber risk.

That result is the load-bearing safety finding of the release. The strongest fine-tuning attack OpenAI could mount, using its own considerable infrastructure and expertise, failed to push the model over the threshold the company considers dangerous. The test is notable precisely because OpenAI simulated the attacker rather than waiting for one to appear.

Second: would releasing gpt-oss-120b significantly advance the frontier of biological capabilities in open foundation models? Here again the answer was no. For most of the evaluations, OpenAI found that the default performance of one or more existing open models comes near to matching the adversarially fine-tuned performance of gpt-oss-120b.

That finding reframes the release decision. If models already available in the open ecosystem perform comparably on biological-capability evaluations, then releasing gpt-oss-120b does not hand the world a capability it lacks. The frontier question — would this specific release move the state of the open-model art in a dangerous direction — is distinct from the question of whether the model is individually capable, and OpenAI addressed it directly.

The stakes for the open-weights debate

The release lands in the middle of an unresolved industry argument. Critics of open-weight releases argue that freely downloadable models cannot be recalled, patched, or restricted once released — the exact concern OpenAI's own model card articulates about fine-tuning attacks. Advocates counter that open models enable independent safety research, customization, competition, and transparency, including the full chain-of-thought visibility that gpt-oss provides.

OpenAI's model card straddles the argument with unusual candor. It states the risk plainly: released weights cannot be revoked, and attackers can fine-tune them without OpenAI's knowledge or intervention. It then presents empirical evidence that this particular model, at this particular capability level, stays below the thresholds the company tracks — even under adversarial fine-tuning conditions.

The message to policymakers and enterprise buyers is that the safety case rests on measurement, not on control. Proprietary models are safe, in part, because OpenAI can intervene post-deployment. Open models must be safe by construction — capable enough to be useful, not capable enough to be dangerous, and evaluated under assumed attack.

For enterprises evaluating gpt-oss for production, the practical checklist follows directly from the documentation. The models support the Responses API, Structured Outputs, tool use, and adjustable reasoning effort, which covers the core requirements of agentic deployments. They follow OpenAI's safety policies by default. But the system-level protections of OpenAI's hosted products do not travel with the weights, and deployments in sensitive contexts will need their own safeguards layered on top.

OpenAI frames the launch as part of a broader commitment, stating that it is "reaffirming its commitment to advancing beneficial AI and raising safety standards across the ecosystem." The operative question for the coming months is whether the ecosystem reciprocates — whether the wide range of stakeholders now building on these weights implements the safeguards that OpenAI's hosted systems would have provided automatically. The model card has defined the shared responsibility. The deployments will test it.

Source: OpenAI News

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

139 articles

Related articles

  1. OpenAI Releases Open-Weight Moderation Models That Reason From Policy
  2. OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
  3. OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard
  4. OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
  5. OpenAI swaps GPT-4o for o3 inside Operator, leaves API unchanged

« Previous articleNext article »