OpenAI Releases Privacy Filter, an Open-Weight PII Redaction Model
OpenAI's new open-weight Privacy Filter detects and redacts PII with 97.43% F1 on a corrected benchmark, runs locally, and ships under Apache 2.0 for commercial use.

Updated
Why it matters
- OpenAI Privacy Filter is an open-weight PII detection and redaction model released under Apache 2.0 on Hugging Face and GitHub
- The 1.5B-parameter model (50M active) achieves 97.43% F1 on a corrected version of the PII-Masking-300k benchmark and supports 128,000 tokens of context
- Fine-tuning on a small amount of domain data raised F1 from 54% to 96% on OpenAI's domain-adaptation benchmark
OpenAI has released Privacy Filter, an open-weight model for detecting and redacting personally identifiable information in text, available today under the Apache 2.0 license on Hugging Face and GitHub. The company says the model achieves state-of-the-art performance on the PII-Masking-300k benchmark — when corrected for annotation issues OpenAI itself identified during evaluation — posting an F1 score of 97.43%.
The release matters for a practical reason: PII detection is plumbing for the AI era. Every company running models over customer logs, support chats, documents, or codebases needs to strip personal data before it reaches training sets, indexes, or third-party APIs. OpenAI is handing developers a piece of that infrastructure for free, in a form they can inspect, fine-tune, and run entirely on their own hardware.
A small model built for one job
Privacy Filter is not a chatbot. It is a bidirectional token-classification model with span decoding, built from an autoregressive pretrained checkpoint and adapted into a token classifier over a fixed taxonomy of privacy labels. Instead of generating text token by token, it labels an input sequence in one pass and decodes coherent spans using a constrained Viterbi procedure.
The released model has 1.5B total parameters with 50M active parameters, and supports up to 128,000 tokens of context. OpenAI highlights four architectural properties aimed at production use: all tokens are labeled in a single forward pass; a language prior lets the model detect PII spans from surrounding context; the long-context window handles lengthy documents; and developers can tune operating points to trade off recall and precision depending on their workflow.
The parameter count matters as much as the benchmark numbers. Because the model is small enough to run locally, data that has yet to be filtered can stay on device rather than being sent to a server for de-identification. That reduces exposure risk at exactly the moment when the data is most sensitive — before it has been cleaned.
Context over regex
Traditional PII detection tools lean on deterministic rules for formats like phone numbers and email addresses. They work for narrow cases but often miss subtler personal information and struggle with context. Privacy Filter is built with deeper language and context awareness, combining strong language understanding with a privacy-specific labeling system.
The practical payoff: the model can better distinguish information that should be preserved because it is public from information that should be masked or redacted because it relates to a private individual. That distinction — a named public figure versus a private citizen, a business address versus a home address — is where rule-based tools routinely fail.
The model predicts spans across eight categories: private_person, private_address, private_email, private_phone, private_url, private_date, account_number, and secret. The account_number category covers a wide variety of account numbers, including credit card numbers and bank account numbers. The secret category targets passwords and API keys. Labels are decoded with BIOES span tags, which OpenAI says produces cleaner and more coherent masking boundaries.
Benchmark results and a caveat worth noting
On the PII-Masking-300k benchmark, Privacy Filter scores 96% F1, with 94.04% precision and 98.04% recall. On a corrected version of the benchmark that accounts for dataset annotation issues identified during review, F1 rises to 97.43%, with 96.79% precision and 98.08% recall.
The framing deserves attention. OpenAI achieved its headline state-of-the-art claim on a version of the benchmark it corrected itself, after finding annotation problems in the public dataset. The company is transparent about this, but readers comparing numbers across vendors should note which version of the benchmark any given result refers to.
Fine-tuning works efficiently. OpenAI reports that fine-tuning on a small amount of data improves accuracy on domain-specific tasks, raising F1 from 54% to 96%, approaching saturation on the domain-adaptation benchmark the company evaluated. OpenAI also tested the model on harder, context-sensitive synthetic and chat-style evaluations, and the model card reports targeted evaluation on secret detection in codebases plus stress tests across multilingual, adversarial, and context-dependent examples.
How it was built
The development process had three stages. First, OpenAI built a privacy taxonomy defining the span types the model should detect: personal identifiers, contact details, addresses, private dates, account numbers such as credit and banking information, and secrets such as API keys and passwords.
Second, the company converted a pretrained language model into a bidirectional token classifier by replacing the language modeling head with a token-classification head and post-training it with a supervised classification objective.
Third, training ran on a mixture of publicly available and synthetic data designed to capture both realistic text and difficult privacy patterns. Where labels in the public data were incomplete, OpenAI used model-assisted annotation and review to improve coverage. Synthetic examples increased diversity across formats, contexts, and privacy subtypes.
OpenAI already runs a fine-tuned version of Privacy Filter in its own privacy-preserving workflows. The company says it built the model because it believed the latest AI capabilities could raise the standard for privacy beyond what was already on the market.
What it is not
OpenAI is explicit about the model's limits, and the caveats are substantive. Privacy Filter is not an anonymization tool, not a compliance certification, and not a substitute for policy review in high-stakes settings. The company describes it as one component in a broader privacy-by-design system.
Its behavior reflects the label taxonomy and decision boundaries it was trained on. Different organizations may want different detection or masking policies, and those policies may require in-domain evaluation or further fine-tuning. Performance may vary across languages, scripts, naming conventions, and domains that differ from the training distribution.
The model can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over- or under-redact entities when context is limited, especially in short sequences. In high-sensitivity domains such as legal, medical, and financial workflows, OpenAI says human review and domain-specific evaluation and fine-tuning remain important.
Release strategy and context
The Apache 2.0 license permits experimentation, customization, and commercial deployment, and the model can be fine-tuned for different data distributions and privacy policies. OpenAI is also publishing documentation covering the model architecture, label taxonomy, decoding controls, intended use cases, evaluation setup, and known limitations, so teams can understand both what the model does well and where it should be used carefully.
The release fits a pattern. OpenAI frames Privacy Filter as part of a broader effort to support a more resilient software ecosystem, citing prior tools and models that make privacy and security protections easier to implement from the start — a reference to its Codex security research preview and its work scaling trusted access for cyber defense. OpenAI has previously released open-weight models such as GPT-OSS, but Privacy Filter is a different bet: not a general-purpose model, but a narrow, efficient tool aimed at a compliance and security problem every AI-deploying organization faces.
The strategic logic is straightforward. "Our goal is for models to learn about the world, not about private individuals," OpenAI writes. "Privacy Filter helps make that possible."
OpenAI positions the release as a preview intended to gather feedback from the research and privacy community and to iterate further on model performance. It also frames the release as evidence of a direction the company believes is important: small, efficient models with frontier capability in narrowly defined tasks that matter for real-world AI systems, released as infrastructure that is easier to inspect, run, adapt, and improve.
For developers, the immediate question is deployment fit. A locally runnable, fine-tunable, commercially licensed PII detector with 128k-token context slots directly into training, indexing, logging, and review pipelines — and if OpenAI's fine-tuning numbers hold up outside its own benchmarks, expect Privacy Filter to become a default component in privacy toolchains rather quickly.
Original: huggingface.co
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles
Related articles
- OpenAI adds Lockdown Mode and risk labels to ChatGPT
- OpenAI Releases MentalHealthBench, an Expert-Built AI Mental Health Benchmark
- OpenAI agents leaked 53 ChatGPT user images
- OpenAI Launches ChatGPT for Teachers, Free Through June 2028
- OpenAI and Anthropic Investigate Tens of Thousands of AI Agent Hacks