OpenAI ships prompt-based teen safety policies for gpt-oss-safeguard
OpenAI has released prompt-based teen safety policies for developers building on gpt-oss-safeguard, its moderation model for catching age-specific risks in AI systems.

Updated
Why it matters
- OpenAI released prompt-based teen safety policies for developers using gpt-oss-safeguard.
- The policies are designed to help developers moderate age-specific risks in AI systems.
- OpenAI frames the release as helping developers build safer AI experiences for teens.
OpenAI has released prompt-based teen safety policies for developers using gpt-oss-safeguard, giving teams that build on the company's open-weight safety model a concrete framework for moderating age-specific risks in AI systems.
The policies take the form of prompts rather than hard-coded rules. Developers can adapt them directly, pairing the guidance with gpt-oss-safeguard to classify and filter content that poses particular risks to teenagers. OpenAI describes the release as a way to help developers build safer AI experiences for teens.
The stakes for the industry are straightforward. Regulators and lawmakers in the United States and Europe have pressed AI companies and app developers over minors' safety, and platforms deploying chatbots to younger users face growing pressure to demonstrate that moderation systems work. Open-sourced safety tooling with explicit policy prompts gives smaller developers — who often lack their own trust-and-safety teams — a starting template they can audit and modify.
The prompt-based approach also matters technically. Policies expressed as prompts can be inspected, versioned and edited by any developer, unlike opaque filters buried inside a hosted API. That transparency is the point of releasing them alongside gpt-oss-safeguard rather than keeping the moderation logic internal.
OpenAI positions the release as practical support for builders: the policies are designed to help developers moderate age-specific risks in the AI systems they ship, from chatbots to assistants, without each team writing teen-safety classification logic from scratch.
The move signals where OpenAI expects responsibility to sit as its models spread across third-party products. By publishing safety policies as adaptable prompts, the company pushes part of the teen-safety implementation work to developers while supplying the scaffolding to do it.
Source: OpenAI News
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- OpenAI Details Mental Health Safety Push Across ChatGPT
- OpenAI adds Lockdown Mode and risk labels to ChatGPT
- OpenAI Publishes Teen Safety Blueprint as a Policy Template for Youth AI Rules
- OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
- OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard