Safety & Security

OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard

OpenAI's gpt-oss-safeguard-120b and -20b classify content against developer-written policies at inference time, beating gpt-5-thinking on multi-policy accuracy under Apache 2.0.

Introducing gpt-oss-safeguard
Introducing gpt-oss-safeguardAI-generated
By Rebecca Stone6 min read

Updated

Why it matters

  • OpenAI released gpt-oss-safeguard-120b and gpt-oss-safeguard-20b under Apache 2.0 on Hugging Face as a research preview.
  • The models classify content against developer-supplied policies at inference time, with reviewable chain-of-thought reasoning, eliminating the need to retrain classifiers when policies change.
  • Both gpt-oss-safeguard models outperformed gpt-5-thinking on OpenAI's internal multi-policy accuracy evaluation; safety reasoning consumed up to 16% of total compute in some recent OpenAI launches.

OpenAI has released a research preview of gpt-oss-safeguard, open-weight reasoning models for safety classification, in two sizes: gpt-oss-safeguard-120b and gpt-oss-safeguard-20b. Both are available now on Hugging Face under the permissive Apache 2.0 license, meaning anyone can use, modify, and deploy them freely.

The release matters because it targets one of the most persistent bottlenecks in content moderation: traditional safety classifiers require thousands of manually labeled examples, and every policy change forces a full retraining cycle. gpt-oss-safeguard takes a different route. The models interpret a developer-provided policy directly at inference time, classifying user messages, completions, and full chats according to that policy. The developer supplies the policy and decides what it says, so outputs are tailored to their specific use case.

The approach is a fine-tuned version of OpenAI's gpt-oss open models. The model takes two inputs at once — a policy and the content to classify under that policy — and outputs a conclusion about where the content falls, along with its chain-of-thought reasoning. Developers can review that reasoning to understand how the model reached its decision. Because the policy is supplied during inference rather than trained into the weights, developers can iteratively revise policies to improve performance without retraining. OpenAI says it developed this approach for internal use first, and that it is significantly more flexible than the traditional method of training a classifier to indirectly infer a decision boundary from a large number of labeled examples.

The use cases OpenAI describes are concrete. A video gaming forum could write a policy to flag posts discussing in-game cheating. A product reviews site could screen reviews that appear likely to be fake. Beyond safety policies, OpenAI says the models can label content in other ways relevant to specific products and platforms.

Where reasoning-based classification works best

OpenAI identifies four situations where the reasoning approach performs especially well: when the potential harm is emerging or evolving and policies need to adapt quickly; when the domain is highly nuanced and difficult for smaller classifiers; when developers lack enough samples to train a high-quality classifier for each risk on their platform; and when latency matters less than producing high-quality, explainable labels.

That last caveat is significant. Traditional classifiers, including those in OpenAI's own Moderation API, offer high performance with low latency and operating cost. Their weakness is cost and time: gathering sufficient training examples is expensive, and updating a policy requires retraining the classifier. In the traditional setup, the classifier never actually sees the safety policy — it infers one indirectly by finding patterns in labeled safe and unsafe content.

From internal tool to open release

gpt-oss-safeguard is an open-weight implementation of an internal OpenAI tool called Safety Reasoner. OpenAI built it through reinforcement fine-tuning on policy labeling tasks, rewarding the model for mirroring correct judgments from human experts, which taught the model to reason about how a policy leads to a judgment.

Safety Reasoner now lets OpenAI dynamically update safety policies in production faster than retraining a classifier would take. The company describes it as a key tool for iterative deployment: new models launch under stricter policies with relatively large amounts of compute applied to careful policy enforcement, then policies loosen as understanding of real-world risks improves. In some recent launches, the fraction of total compute devoted to safety reasoning has ranged as high as 16%.

The tool has become a core component of OpenAI's safety stack. For image generation and Sora 2, it performs dynamic, step-wise evaluations to identify and block unsafe generations in real time. In domains such as biology and self-harm, OpenAI runs small, fast Moderation API-class models as high-recall filters to determine which content falls in scope, then applies Safety Reasoner to review it. Safety Reasoner also classifies model outputs against a detailed taxonomy across systems like GPT-5 and ChatGPT Agent. This connects to what OpenAI calls deliberative alignment — its reasoning models now learn safety policies directly and reason about what is safe, which the company says makes them safer on several axes than their non-reasoning predecessors.

Benchmark results

OpenAI evaluated the models on internal and external sets. In the internal evaluation, the model received multiple policies simultaneously and was counted as accurate only if it exactly matched golden-set labels across all included policies — a strict standard. Both gpt-oss-safeguard models and the internal Safety Reasoner outperformed gpt-5-thinking and the gpt-oss open models on multi-policy accuracy. OpenAI calls the gpt-oss-safeguard models beating gpt-5-thinking "especially surprising given the former models' small size."

On the moderation dataset OpenAI released with its 2022 research paper, evaluated under internal safety policies, gpt-oss-safeguard slightly outperformed all other tested models, including Safety Reasoner and gpt-5-thinking — though the gap versus Safety Reasoner was not statistically significant. On ToxicChat, a public benchmark based on user queries to an open-source chatbot, Safety Reasoner again beat gpt-5-thinking, while both of those marginally outperformed gpt-oss-safeguard-120b and gpt-oss-safeguard-20b. OpenAI still expects the models' small size to make them preferable for that type of task.

Stated limitations

OpenAI is explicit about two limitations. First, classifiers trained on tens of thousands of high-quality labeled samples can still outperform gpt-oss-safeguard when it reasons directly from a policy — for higher performance on more complex risks, training a dedicated classifier may remain the better choice. Second, gpt-oss-safeguard is time and compute-intensive, making it hard to scale across all platform content. Internally, OpenAI mitigates this by using smaller, faster classifiers to decide which content to assess, and by running Safety Reasoner asynchronously in some circumstances to preserve a low-latency user experience while retaining the ability to intervene on unsafe content.

Community rollout

This is OpenAI's first set of open safety models built with the community. OpenAI developed the release over months with ROOST to identify developer needs, test the model, and produce documentation, and iterated with trust and safety specialists at SafetyKit, ROOST, Tomoro, and Discord during early testing. ROOST is also launching a model community today to explore open AI models for protecting online spaces.

ROOST CTO Vinay Rao offered a pointed framing of the release: "gpt-oss-safeguard is the first open source reasoning model with a 'bring your own policies and definitions of harm' design. Organizations deserve to freely study, modify and use critical safety technologies and be able to innovate. In our testing, it was skillful at understanding different policies, explaining its reasoning, and showing nuance in applying the policies, which we believe will be beneficial to builders and safety teams."

OpenAI published a technical report alongside the release detailing the preview model's safety performance, and says it will keep iterating through the ROOST Model Community, which brings safety practitioners and researchers together to share implementation practices, evaluation outcomes, and model feedback. For teams that lack the labeled data to train dedicated classifiers — or that need policies to shift as fast as the harms they target — the models are downloadable from Hugging Face today.

Original: huggingface.co

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
  2. OpenAI Releases Open-Weight Moderation Models That Reason From Policy
  3. OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
  4. OpenAI Updates GPT-5 System Card With New Mental Health Evals
  5. OpenAI Releases GPT-5.2 With Safety Approach Carried Over

« Previous articleNext article »