Models

OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context

OpenAI's GPT‑5.4 brings native computer use, 1M-token context and tool search to ChatGPT, the API and Codex, beating human performance on OSWorld at 75.0% while raising API prices.

Introducing GPT-5.4
Introducing GPT-5.4AI-generated
By Elena Vasquez7 min read

Updated

Why it matters

  • GPT‑5.4 scores 75.0% on OSWorld-Verified, surpassing the cited human performance of 72.4% and GPT‑5.2's 47.3%, as OpenAI's first general-purpose model with native computer-use capabilities.
  • API pricing rises to $2.50/M input and $15/M output tokens for gpt-5.4 (from $1.75/$14 for gpt‑5.2); gpt-5.4-pro costs $30/M input and $180/M output.
  • Tool search cuts token usage by 47% at equal accuracy on 250 MCP Atlas tasks with 36 MCP servers enabled; GPT‑5.4 Pro sets a BrowseComp state of the art at 89.3%.

OpenAI has released GPT‑5.4 across ChatGPT, the API and Codex, calling it the company's "most capable and efficient frontier model for professional work" and the first general-purpose model it has shipped with native computer-use capabilities. A higher-performance variant, GPT‑5.4 Pro, arrives simultaneously in ChatGPT and the API.

The release matters beyond benchmark tables. It merges the coding strength of GPT‑5.3‑Codex into OpenAI's mainline reasoning model, adds agents that can operate desktop software, and raises per-token API prices — a signal that OpenAI believes frontier agentic capability commands a premium in an increasingly competitive market for autonomous work tools.

Computer use beats humans on OSWorld

The headline technical result: GPT‑5.4 scores a state-of-the-art 75.0% success rate on OSWorld-Verified, which measures a model's ability to navigate a desktop environment through screenshots and keyboard/mouse actions. That exceeds GPT‑5.2's 47.3% and surpasses the human performance figure of 72.4% that OpenAI cites.

The model writes code to operate computers via libraries like Playwright and can issue mouse and keyboard commands in response to screenshots. Developers can steer behavior through developer messages and configure custom confirmation policies to match different levels of risk tolerance.

Browser-based results are similarly strong. GPT‑5.4 achieves 67.3% on WebArena-Verified using both DOM- and screenshot-driven interaction (GPT‑5.2: 65.4%), and 92.8% on Online-Mind2Web using screenshot observations alone, ahead of ChatGPT Atlas's Agent Mode at 70.9%.

GPT‑5.4 supports up to 1M tokens of context in Codex, where it ships with experimental support for the full window. Requests exceeding the standard 272K context window count against usage limits at 2x the normal rate.

Knowledge work: 83.0% on GDPval

On GDPval, OpenAI's benchmark testing agent performance on well-specified knowledge work across 44 occupations from the top nine industries contributing to U.S. GDP, GPT‑5.4 matches or exceeds industry professionals in 83.0% of comparisons, up from 70.9% for GPT‑5.2.

OpenAI placed particular emphasis on spreadsheets, presentations and documents. On an internal benchmark of spreadsheet modeling tasks that a junior investment banking analyst might do, GPT‑5.4 scores a mean of 87.3%, versus 68.4% for GPT‑5.2. Human raters preferred presentations from GPT‑5.4 68.0% of the time over those from GPT‑5.2, citing stronger aesthetics, greater visual variety and more effective use of image generation.

Enterprise customers get a new ChatGPT for Excel add-in, launched the same day, and OpenAI has updated its spreadsheet and presentation skills available in Codex and the API.

OpenAI also claims factual improvements. On a set of de-identified prompts where users flagged factual errors, GPT‑5.4's individual claims are 33% less likely to be false and its full responses are 18% less likely to contain any errors, relative to GPT‑5.2.

Coding and tool search

GPT‑5.4 matches or outperforms GPT‑5.3‑Codex on SWE-Bench Pro (Public) at 57.7% versus 56.8%, while delivering lower latency across reasoning efforts. A new /fast mode in Codex delivers up to 1.5x faster token velocity with the same model, according to OpenAI, and developers can reach equivalent speeds through the API's priority processing.

The API introduces tool search, a structural change to how models handle tool definitions. Previously, all tool definitions sat in the prompt upfront, potentially adding tens of thousands of tokens to every request. GPT‑5.4 instead receives a lightweight list of available tools and looks up definitions only when needed.

The efficiency gains are measurable. OpenAI evaluated 250 tasks from Scale's MCP Atlas benchmark with all 36 MCP servers enabled in two modes — every function exposed directly, and all servers behind tool search. The tool-search configuration cut total token usage by 47% while achieving the same accuracy. On the full MCP Atlas benchmark, GPT‑5.4 scores 67.2% versus 60.6% for GPT‑5.2.

On Toolathlon, which tests multi-step tool workflows — for example, reading emails, extracting attachments, uploading them, grading them and recording results in a spreadsheet — GPT‑5.4 scores 54.6%, up from 45.7%. On BrowseComp, which measures persistent web browsing for hard-to-locate information, GPT‑5.4 jumps 17 points absolute over GPT‑5.2 to 82.7%, while GPT‑5.4 Pro sets a new state of the art at 89.3%.

OpenAI notes it used a search blocklist excluding websites containing benchmark answers to prevent contamination, and that GPT‑5.4 was tested with a longer, updated blocklist at a later date than GPT‑5.2.

Steerability and vision

In ChatGPT, GPT‑5.4 Thinking now presents an upfront plan of its thinking on longer, complex queries, letting users adjust course mid-response rather than starting over. The feature is live on chatgpt.com and Android, with iOS to follow.

Vision improvements underpin the computer-use gains. GPT‑5.4 scores 81.2% on MMMU-Pro without tools (GPT‑5.2: 79.5%). On document parsing, measured by OmniDocBench normalized edit distance with reasoning disabled, error drops to 0.109 from 0.140. A new original image input detail level supports full-fidelity perception up to 10.24M total pixels or a 6000-pixel maximum dimension; the high level now supports up to 2.56M total pixels or 2048 pixels. Early testing with API users showed strong gains in localization, image understanding and click accuracy, OpenAI reports.

Safety and pricing

OpenAI classifies GPT‑5.4 as High cyber capability under its Preparedness Framework, deploying it with an expanded cyber safety stack including monitoring systems, trusted access controls and asynchronous blocking for higher-risk requests on Zero Data Retention surfaces. The company acknowledges some false positives may occur as classifiers improve.

The company also released an open-source evaluation, CoT controllability, measuring whether models can deliberately obfuscate their reasoning to evade monitoring. OpenAI reports GPT‑5.4 Thinking's ability to control its chain of thought is low — a positive property, the company argues, because it suggests the model cannot hide its reasoning and CoT monitoring remains effective.

Pricing rises from GPT‑5.2. In the API, gpt-5.4 costs $2.50 per million input tokens ($0.25 cached) and $15 per million output tokens, against $1.75/$0.175/$14 for gpt-5.2. GPT‑5.4 Pro is priced at $30 per million input tokens and $180 per million output tokens, versus $21/$168 for gpt-5.2-pro. Batch and Flex pricing is half the standard rate; priority processing costs twice the standard rate. OpenAI says GPT‑5.4 is its most token-efficient reasoning model yet, using significantly fewer tokens to solve problems than GPT‑5.2, which it says offsets the higher per-token cost on many tasks.

GPT‑5.4 Thinking is available now for ChatGPT Plus, Team and Pro users, replacing GPT‑5.2 Thinking, which remains in the model picker for three months and retires on June 5, 2026. Enterprise and Edu admins can enable early access. GPT‑5.4 Pro is limited to Pro and Enterprise plans.

Notable academic results include 73.3% on ARC-AGI-2 (Verified) for GPT‑5.4 and 83.3% for GPT‑5.4 Pro — a sharp jump from GPT‑5.2's 52.9% — and 50.0% for GPT‑5.4 Pro on FrontierMath Tiers 1–3. One regression stands out: on Terminal-Bench 2.0, GPT‑5.4 scores 75.1%, below GPT‑5.3‑Codex's 77.3%. Long-context retrieval beyond 256K remains difficult, with Graphwalks BFS accuracy at 21.4% and 8-needle MRCR accuracy at 36.6% in the 512K–1M range.

OpenAI frames the version jump to 5.4 as a deliberate simplification: it is the first mainline reasoning model to absorb the frontier coding capabilities of the Codex line, collapsing the choice between models in Codex. The company says Instant and Thinking models will continue to evolve at different speeds — a hint that the gap between fast, cheap models and frontier reasoning models will widen rather than close.

Original: chatgpt.com

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI ships GPT-5.3-Codex, its first self-built coding model
  2. OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
  3. OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
  4. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  5. OpenAI says GPT-5.6 Sol cut its own serving costs by 20 percent

« Previous articleNext article »