Models

OpenAI's GPT-5 System Card Reveals Router Architecture

OpenAI's GPT-5 system card details a unified system: a real-time router picks between fast and thinking models, with GPT-4o and o3 successors across six variants and a High bio-risk rating.

GPT-5 System Card
GPT-5 System Cardpenelope waits / Openverse
By Elena Vasquez6 min read

Updated

Why it matters

  • GPT-5 uses a real-time router trained on user switching, preference rates, and measured correctness to choose between fast (gpt-5-main) and thinking (gpt-5-thinking) models.
  • The GPT-5 lineup replaces prior models: gpt-5-main succeeds GPT-4o, gpt-5-thinking succeeds o3, gpt-5-thinking-pro succeeds o3 Pro, alongside mini and nano variants.
  • OpenAI classifies gpt-5-thinking as High capability in the Biological and Chemical domain under its Preparedness Framework as a precautionary measure, despite lacking definitive evidence it meets the threshold.

OpenAI's GPT-5 system card, published alongside the model's release, describes a unified architecture built on a real-time router that decides which of two model families answers each query — and the document discloses the full lineage of how GPT-5's variants replace the company's previous flagship models.

The system card lays out six named models. The fast, high-throughput models carry the labels gpt-5-main and gpt-5-main-mini; the thinking models are gpt-5-thinking and gpt-5-thinking-mini. The API adds an even smaller developer-oriented variant, gpt-5-thinking-nano, while ChatGPT exposes gpt-5-thinking through a setting that uses parallel test time compute, which OpenAI calls gpt-5-thinking-pro.

The succession table in the system card makes the mapping explicit: gpt-5-main replaces GPT-4o, gpt-5-main-mini replaces GPT-4o-mini, gpt-5-thinking replaces OpenAI o3, gpt-5-thinking-mini replaces o4-mini, gpt-5-thinking-nano replaces GPT-4.1-nano, and gpt-5-thinking-pro replaces OpenAI o3 Pro. In other words, GPT-5 folds the previously separate GPT-4o product line and the o-series reasoning models into a single system, a structural change that ends the split between OpenAI's chat and reasoning branches.

The router itself is the central piece of the design. It decides which model handles a query based on, in OpenAI's words, "conversation type, complexity, tool needs, and explicit intent (for example, if you say 'think hard about this' in the prompt)." A user who explicitly asks for deeper reasoning can force the thinking model; otherwise the system decides on its own.

The routing is not static. "The router is continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness, improving over time," the system card states. That means the dispatch decisions users see today may shift as OpenAI trains the router on accumulated behavioral data — a design choice with cost, latency, and quality implications for every downstream application.

There is also a fallback mechanism: "Once usage limits are reached, a mini version of each model handles remaining queries." Users who exhaust their allocation of the full-size models will not hit a hard wall; they will slide down to the mini tier for the remainder of the query window.

OpenAI says this architecture is temporary. "In the near future, we plan to integrate these capabilities into a single model," the system card reads, signaling that the two-family-plus-router design is a transitional stage toward one model that internally adapts its compute to the task.

Performance and behavior claims

On capabilities, OpenAI frames GPT-5's gains in operational terms rather than raw benchmark scores. "The GPT-5 system not only outperforms previous models on benchmarks and answers questions more quickly, but—more importantly—is more useful for real-world queries," the system card states. The document highlights work on three failure modes that have dogged large language models: "We've made significant advances in reducing hallucinations, improving instruction following, and minimizing sycophancy."

The company also names the three domains where it concentrated effort: "leveled up GPT-5's performance in three of ChatGPT's most common uses: writing, coding, and health." Those three categories map to where ChatGPT actually gets used at scale, which makes them the surface where perceived model quality will be judged by the largest number of users.

Every GPT-5 model ships with a new safety layer. "All of the GPT-5 models additionally feature safe-completions, our latest approach to safety training to prevent disallowed content," the system card says. Unlike safety mechanisms applied only at the system-prompt or moderation layer, safe-completions operates in training, applying across the full model lineup including the nano and mini variants that developers access through the API.

The biological risk classification

The most consequential safety decision in the document concerns biological and chemical threats. OpenAI has assigned gpt-5-thinking a High capability rating in the Biological and Chemical domain under its Preparedness Framework, which triggers the framework's associated safeguards.

The rating is precautionary rather than evidence-driven. "While we do not have definitive evidence that this model could meaningfully help a novice to create severe biological harm—our defined threshold for High capability—we have chosen to take a precautionary approach," the system card states. High is OpenAI's defined threshold in the Preparedness Framework for models that could plausibly assist in severe biological harm; the framework document governs what mitigations must be activated before deployment.

The phrasing matters. OpenAI is saying the model sits near enough to the threshold that it is treating the risk as if it had crossed it. This mirrors the approach the company took with ChatGPT agent, which the system card references directly: "Similarly to ChatGPT agent, we have decided to treat gpt-5-thinking as High capability in the Biological and Chemical domain under our Preparedness Framework, activating the associated safeguards."

The decision applies to the thinking model specifically, not the entire lineup — a recognition that extended reasoning is what elevates uplift risk. The fast models, gpt-5-main and its mini, are not given the same classification in the card.

Scope of the document

The system card notes its own coverage limits: "This system card focuses primarily on gpt-5-thinking and gpt-5-main, while evaluations for other models are available in the appendix." Detailed evaluations for the mini, nano, and pro variants live outside the main body of the document.

Why the architecture matters

The router-based design answers a question that has hung over OpenAI's product strategy since the o-series launch: how to reconcile a fast general-purpose model with slower, more expensive reasoning models when users do not want to choose. GPT-5's answer is to remove the choice. The system assesses each conversation and routes it, with an escape hatch for explicit user intent.

The continuously trained router changes the economics. OpenAI is optimizing dispatch against "when users switch models, preference rates for responses, and measured correctness" — signals that tie routing decisions directly to user satisfaction data. For API developers, the named endpoints (gpt-5-thinking, gpt-5-thinking-mini, gpt-5-thinking-nano) preserve direct access to reasoning compute at three price-performance points, while the router governs the ChatGPT experience.

The pledge to consolidate matters most for what comes next. "In the near future, we plan to integrate these capabilities into a single model," OpenAI writes, which would make the router unnecessary once a single model can modulate its own reasoning depth — the same adaptive-compute direction the rest of the frontier-lab field is pursuing. Until then, GPT-5 users will be interacting with a system that quietly picks its own brain for them, and the mini-tier fallback after usage limits will define what heavy users actually experience day to day.

Original: cdn.openai.com

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

137 articles

Related articles

  1. OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
  2. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  3. OpenAI Launches ChatGPT Go Worldwide With GPT-5.2 Instant Access
  4. OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context
  5. OpenAI and Anthropic Ship Cheaper, Faster Models Same Day

« Previous articleNext article »