Safety & Security

OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models

OpenAI's gpt-oss-safeguard-120b and -20b classify content against developer-written policies at inference time, beating gpt-5-thinking on multi-policy accuracy, under Apache 2.0.

By Marcus Bennett7 min read

Updated

Why it matters

  • OpenAI released gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, open-weight safety classification models under Apache 2.0, downloadable from Hugging Face.
  • The models interpret developer-provided policies at inference time via chain-of-thought reasoning, outperforming gpt-5-thinking on OpenAI's internal multi-policy accuracy evaluation.
  • The release builds on OpenAI's internal Safety Reasoner tool, which in some recent launches consumed up to 16% of total compute devoted to safety reasoning.

OpenAI has released a research preview of gpt-oss-safeguard, a pair of open-weight reasoning models for safety classification, available in two sizes: gpt-oss-safeguard-120b and gpt-oss-safeguard-20b. Both are fine-tuned versions of OpenAI's gpt-oss open models, carry the permissive Apache 2.0 license, and can be downloaded today from Hugging Face.

The release matters because it targets one of the most persistent bottlenecks in trust and safety work: building classifiers for harms that evolve faster than labeled datasets can be collected. If the approach performs as OpenAI reports, small safety teams without the resources to curate tens of thousands of labeled examples get a workable alternative.

Policies at inference time, not in training

The gpt-oss-safeguard models use reasoning to directly interpret a developer-provided policy at inference time, classifying user messages, completions, and full chats according to the developer's needs. The developer decides what policy to use. The model produces a chain-of-thought the developer can review to understand how it reached its decision.

Because the policy is supplied during inference rather than trained into the model, developers can iteratively revise policies without retraining. OpenAI states this approach, which it initially developed for internal use, is significantly more flexible than the traditional method of training a classifier to indirectly infer a decision boundary from a large number of labeled examples.

The model takes two inputs at once — a policy and the content to classify under that policy — and outputs a conclusion about where the content falls, along with its reasoning. Developers decide how, if at all, to use those conclusions in their own safety pipelines.

Use cases vary widely. A video gaming discussion forum might build a policy to classify posts discussing cheating in the game. A product reviews site might use its own policy to screen reviews that appear likely to be fake.

OpenAI says the reasoning-based approach performs especially well when the potential harm is emerging or evolving and policies need to adapt quickly, when the domain is highly nuanced and difficult for smaller classifiers to handle, when developers lack enough samples to train a high-quality classifier for each risk, and when latency matters less than producing high-quality, explainable labels.

How it differs from the Moderation API

Traditional safety classifiers, such as those in OpenAI's Moderation API, are developed by manually curating thousands of examples of safe and unsafe content under pre-defined safety policies. The classifier never actually sees the safety policy. It attempts to infer the underlying policy from patterns in the labeled data.

That approach delivers high performance with low latency and operating cost. But gathering sufficient training examples is time-consuming and costly, and changing the policy requires retraining the classifier.

gpt-oss-safeguard works differently. Its reasoning capabilities let developers apply any policy, including ones they write themselves or draw from other sources, and reasoning helps the models generalize over newly written policies. Beyond safety policies, OpenAI notes the models can label content in other ways important to specific products and platforms.

The internal lineage: Safety Reasoner

gpt-oss-safeguard is an open-weight implementation of an internal OpenAI tool called Safety Reasoner. OpenAI began with reinforcement fine-tuning on policy labelling tasks, rewarding the model for mirroring correct judgments from human experts. That taught the model to reason about how a policy leads to a judgment.

Today, Safety Reasoner lets OpenAI dynamically update its safety policies in production faster than retraining a classifier would allow. The company describes this as a key tool for iterative deployment: it starts with stricter policies and relatively large amounts of compute when launching new models, then adjusts policies as its understanding of production risks improves. In some recent launches, the fraction of total compute devoted to safety reasoning has ranged as high as 16%.

Safety Reasoner now serves as a core component of OpenAI's safety stack. For image generation and Sora 2, it performs dynamic, step-wise evaluations of outputs to identify and block unsafe generations in real time. In domains such as biology and self-harm, OpenAI runs Moderation API-style models as small, fast, high-recall classifiers to determine which content falls within a domain of interest, then uses Safety Reasoner to review that content. Safety Reasoner also classifies model outputs against a detailed taxonomy across systems like GPT-5 and ChatGPT Agent.

The approach builds on what OpenAI calls deliberative alignment, in which its reasoning models learn safety policies directly and reason about what is safe. OpenAI says this significantly improves on earlier safety training methods and makes its reasoning models safer on several axes than their non-reasoning predecessors, even as capabilities increase.

Benchmark results

OpenAI evaluated the models on both internal and external sets.

In the internal evaluation, multiple policies were provided simultaneously at inference time, and the model was counted as accurate only if it exactly matched the golden set labels for all included policies. OpenAI reports that gpt-oss-safeguard and its internal Safety Reasoner outperform gpt-5-thinking and the gpt-oss open models on multi-policy accuracy. The company calls the gpt-oss-safeguard models beating gpt-5-thinking especially surprising given their small size.

On the moderation dataset OpenAI released with its 2022 research paper, evaluated using internal safety policies with one policy adapted for subsequent changes, gpt-oss-safeguard slightly outperformed all other tested models, including internal Safety Reasoner and gpt-5-thinking — though OpenAI notes the difference between Safety Reasoner and gpt-oss-safeguard is not statistically significant.

On ToxicChat, a public benchmark based on user queries to an open-source chatbot, evaluated with a short hand-written prompt adapted from internal policies, Safety Reasoner again outperformed gpt-5-thinking, while both gpt-5-thinking and Safety Reasoner marginally outperformed gpt-oss-safeguard-120b and gpt-oss-safeguard-20b. OpenAI still expects the models' relatively small size to make them preferable for this type of task.

Two limitations worth noting

OpenAI is explicit about where gpt-oss-safeguard falls short. First, classifiers trained on tens of thousands of high-quality labeled samples can still outperform it when it reasons directly from the policy; training a dedicated classifier may be preferred for higher performance on more complex risks. Second, the models are time and compute-intensive, making them hard to scale across all platform content. Internally, OpenAI handles this by using smaller, faster classifiers to decide which content to assess, and in some circumstances running Safety Reasoner asynchronously to keep latency low while retaining the ability to intervene on unsafe content.

Built with the community

OpenAI describes gpt-oss-safeguard as its first set of open safety models built with the community. The company iterated with trust and safety specialists at SafetyKit, ROOST, Tomoro, and Discord during early testing, and worked with ROOST over months to identify developer needs, test the model, and produce documentation. OpenAI also published a technical report detailing the preview model's safety performance.

ROOST CTO Vinay Rao says: "gpt-oss-safeguard is the first open source reasoning model with a 'bring your own policies and definitions of harm' design. Organizations deserve to freely study, modify and use critical safety technologies and be able to innovate. In our testing, it was skillful at understanding different policies, explaining its reasoning, and showing nuance in applying the policies, which we believe will be beneficial to builders and safety teams."

ROOST is also launching a model community, the RMC, to share best practices for implementing open source AI models in safety workflows, including evaluation outcomes and model feedback. OpenAI says it will continue iterating with the community to improve open safety tooling through that partnership — a signal that this research preview is a first step rather than a finished product, with the company explicitly seeking feedback from the research and safety community to improve model performance.

Original: huggingface.co

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

108 articles

Related articles

  1. OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard
  2. OpenAI Releases Open-Weight Moderation Models That Reason From Policy
  3. OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
  4. OpenAI Releases GPT-5.2 With Safety Approach Carried Over
  5. OpenAI ships prompt-based teen safety policies for gpt-oss-safeguard

« Previous articleNext article »