OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
OpenAI's IH-Challenge RL dataset taught GPT-5 Mini-R to prioritize system over developer over user over tool, cutting prompt-injection and jailbreak success while preserving general capability.

Updated
Why it matters
- OpenAI trained an internal model, GPT-5 Mini-R, on the IH-Challenge RL dataset to follow its instruction hierarchy (system > developer > user > tool), with gains generalizing to held-out and adversarial tests.
- Developer-User Conflict robustness improved from 0.83 to 0.95 and TensorTrust (dev-user) from 0.76 to 0.91, while GPQA Diamond held at 0.83 and AIME 2024 ticked up from 0.93 to 0.94.
- OpenAI released the IH-Challenge dataset publicly on Hugging Face to support further instruction-hierarchy research.
OpenAI has trained an internal research model, GPT-5 Mini-R, to more reliably follow its instruction hierarchy — system > developer > user > tool — and reports measurable gains in prompt-injection robustness and safety steerability without significant capability regressions. The company is releasing the training dataset, called IH-Challenge, publicly on Hugging Face to support further research.
The stakes are straightforward. AI systems now receive instructions from multiple sources at once: safety policies in system messages, guidance from developers, requests from users, and content pulled from the web or returned by tools. Many safety and reliability failures — disallowed content requests, attempts to extract private information, prompt-injection attacks embedded in online data — share a single root cause, according to OpenAI: "the model may follow the wrong instruction."
"Getting this right is foundational to safety, security, and reliability," the company writes.
What the instruction hierarchy is
OpenAI's models are trained to prioritize instructions by trust level: system messages outrank developer messages, which outrank user messages, which outrank tool outputs. The model should follow lower-priority instructions only when they do not conflict with higher-priority constraints. These principles are codified in the OpenAI Model Spec.
The consequences are concrete. If a system message contains a safety policy and a user asks the model to violate it, the model should refuse. If a tool output contains malicious instructions, the model should ignore them rather than treat them as commands. In a worked example, OpenAI shows a developer instructing a model to act as a math tutor without giving away answers, while a user pleads for the solution to x² + 2x + 1 = 0. The correctly behaving model follows the developer's higher-priority instruction.
Why training this is hard
Reinforcement learning is a natural fit for teaching the hierarchy: generate conversations with conflicting instructions, prompt the model to respond, and reward correct behavior. But OpenAI identified three pitfalls in applying that recipe naively.
First, instruction-following failures can double as instruction-hierarchy failures — a model might fail to resolve a conflict not because it misunderstands the hierarchy, but because the instructions themselves are too complicated. Second, instruction conflicts can be nuanced and even subjective, and the LLM judges commonly used to assign rewards are themselves fallible. Third, models learn shortcuts that score high reward but are useless in practice. The classic example is over-refusal: a model can maximize safety metrics by refusing even benign requests.
IH-Challenge: simple tasks, objective grading
IH-Challenge, OpenAI's RL training dataset, is designed to sidestep all three pitfalls. The tasks follow three principles: they are "instruction-following-simple," they are "objectively-gradable with a simple Python script," and they contain "no trivial shortcuts that guarantee high reward across all tasks."
Each task is a conversation containing an instruction message from a high-privilege role — for example, "Only answer 'Yes' or 'No'" — followed by an instruction message from a lower-privilege role that tries to get the model to violate the higher-priority constraint. The environments are constructed so that a script can programmatically check whether the model's response satisfies the higher-level constraint. No LLM judge is needed, and the tasks are simple enough that failures clearly indicate hierarchy failures rather than comprehension failures.
Results on academic benchmarks
Training GPT-5 Mini on IH-Challenge produced the internal GPT-5 Mini-R model. OpenAI reports it performs better on instruction-hierarchy benchmarks, that the improvement generalizes to held-out and adversarial tests, and that the model maintains overall usefulness without collapsing into over-refusal.
On academic robustness benchmarks, the gains are uneven but real. Gandalf Password (sys-user) held steady at 0.99, while Gandalf Password (dev-user) rose from 0.98 to 1.00 (+0.02). TensorTrust improved from 0.86 to 0.94 (+0.08) in the sys-user configuration and from 0.76 to 0.91 (+0.15) in the dev-user configuration. RealGuardrails Distractors rose from 0.88 to 0.95 (+0.07) and RealGuardrails Handwritten from 0.82 to 0.89 (+0.07). System IFEval moved from 0.92 to 0.96 (+0.04).
Internal benchmarks show similar movement. TutorJailbreak (sys-user) improved from 0.96 to 0.99 (+0.03), and TutorJailbreak (dev-user) from 0.97 to 0.99 (+0.02). System-User Conflict jumped from 0.84 to 0.95 (+0.11), and Developer-User Conflict from 0.83 to 0.95 (+0.12). System-Developer Conflict stayed flat at 0.86 (+0).
No free lunch, but a small bill
The capability picture is mixed but mostly stable. On the IH-Challenge over-refusal measure, GPT-5 Mini-R scored 1.00 versus the baseline's 0.79 — a +0.21 gain, meaning the trained model almost never refuses benign requests on the tasks it was trained to resolve. TensorTrust over-refusal ticked down slightly from 0.91 to 0.90 (-0.01).
General capability held steady. GPQA Diamond stayed at 0.83, and AIME 2024 rose marginally from 0.93 to 0.94 (+0.01). Two preference metrics declined: Chat WinRate versus o1 fell from 0.71 to 0.66 (-0.05), and Preference Score dropped from 0.46 to 0.40 (-0.06). OpenAI characterizes the result as maintaining overall usefulness without over-refusal collapse, though the preference dips suggest the safety gains carry at least a modest cost in head-to-head human preference.
Why it matters for real-world safety
The improvements show up in two properties OpenAI ties to production safety.
The first is safety steerability. OpenAI evaluated this by adding category-specific safety specifications to system prompts and measuring behavior on its safety Production Benchmarks, described as a set of safety-sensitive conversations representative of ChatGPT in production. With the safety spec present, GPT-5 Mini-R achieved higher refusal and safe completion rates across disallowed categories. Notably, the improvement did not come with a corresponding decrease in helpfulness rate — the model is not simply refusing more overall.
The second is prompt-injection robustness. OpenAI evaluated the trained model on CyberSecEval 2, an academic benchmark, and on an internal prompt-injection benchmark that includes attacks like one previously demonstrated against an older version of ChatGPT Atlas. GPT-5 Mini-R improved over the baseline on both, with what OpenAI calls substantial improvement on the internal static prompt-injection evaluation.
The generalization claim
The most consequential finding may be the generalization result. The IH-Challenge tasks are deliberately simple and objectively gradable, yet the behavior learned from them transferred to "more realistic, often not-objectively-gradable benchmarks," including held-out and adversarial instruction-hierarchy tests and prompt-injection attacks the model never saw in training.
"By directly training models to resolve instruction conflicts correctly on IH-challenge tasks, we get IH improvements that generalize to new attacks and new situations," OpenAI writes.
That matters because the realistic benchmarks are exactly the ones that are hard to train on directly — they involve nuanced, subjective conflicts where LLM judges are unreliable. If simple, programmatically checkable training environments can produce robust behavior on the messy cases, the approach offers a scalable recipe that avoids depending on fallible reward models.
The road ahead
OpenAI frames the work against the trajectory toward agentic AI. "As models become more agentic—calling tools, reading untrusted documents, and taking actions in the world—the ability to consistently prioritize trusted instructions over untrusted ones becomes a core safety property," the company writes.
The release of the IH-Challenge dataset on Hugging Face signals that OpenAI sees instruction-hierarchy training as an area where external researchers can contribute. As models grow more capable and more autonomous, the company argues, strengthening instruction hierarchy delivers reliability plus multiple safety and security gains at once — a foundation that becomes increasingly important, not less, as systems gain the ability to act.
Original: model-spec.openai.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
- OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard
- OpenAI ships GPT-5.4 Thinking with first High-tier cyber mitigations
- OpenAI Releases Open-Weight Moderation Models That Reason From Policy
- OpenAI's GPT-5.1-Codex-Max Flags Coming Cybersecurity Threshold