OpenAI ships GPT-5.3-Codex, its first self-built coding model
OpenAI's GPT-5.3-Codex tops SWE-Bench Pro at 56.8% and jumps to 64.7% on OSWorld. It is OpenAI's first model classified High capability for cyber tasks — and it helped debug its own training.

Updated
Why it matters
- GPT-5.3-Codex scores 56.8% on SWE-Bench Pro and 77.3% on Terminal-Bench 2.0, both industry highs, while running 25% faster than its predecessor.
- OSWorld-Verified computer-use performance jumps from 38.2% (GPT-5.2-Codex) to 64.7%, approaching the ~72% human baseline.
- OpenAI classifies the model as High capability for cybersecurity under its Preparedness Framework, routes risky requests to GPT-5.2, and is committing $10M in API credits to cyber defense.
OpenAI has released GPT-5.3-Codex, a coding and agentic model it calls "the most capable agentic coding model to date" — and the company's first model that "was instrumental in creating itself." The release, announced on OpenAI's blog, combines the frontier coding performance of GPT-5.2-Codex with the reasoning and professional knowledge capabilities of GPT-5.2, while running 25% faster.
The launch matters beyond raw benchmarks. OpenAI is positioning Codex not as a code generator but as a general-purpose computer agent, and it is the first model OpenAI classifies as "High capability" for cybersecurity tasks under its Preparedness Framework — a designation that triggered its most comprehensive cyber safety stack to date.
Benchmark results
GPT-5.3-Codex sets a new industry high on SWE-Bench Pro and Terminal-Bench, and OpenAI reports strong results on OSWorld and GDPval — four benchmarks the company uses to measure coding, agentic and real-world capabilities.
On SWE-Bench Pro, a real-world software engineering evaluation spanning four languages (unlike SWE-bench Verified, which only tests Python) and designed to be more contamination-resistant, GPT-5.3-Codex scores 56.8% at the xhigh compute setting, ahead of GPT-5.2-Codex at 56.4% and GPT-5.2 at 55.6%. On Terminal-Bench 2.0, which measures terminal skills, the new model scores 77.3% — a jump of more than 13 points over GPT-5.2-Codex's 64.0%. OpenAI notes the model achieves these results with fewer tokens than any prior model.
The largest gain appears in computer use. On OSWorld-Verified, an agentic benchmark where models complete productivity tasks in a visual desktop environment using vision, GPT-5.3-Codex scores 64.7%, up from 38.2% for GPT-5.2-Codex and 37.9% for GPT-5.2. Humans score roughly 72%. On GDPval, OpenAI's 2025 evaluation of well-specified knowledge-work tasks across 44 occupations, the model matches GPT-5.2 at 70.9% wins-or-ties. It also scores 77.6% on cybersecurity Capture The Flag challenges, up from 67.4%, and 81.4% on SWE-Lancer IC Diamond.
From coding agent to computer agent
OpenAI frames the release as a shift in scope: "With GPT-5.3-Codex, Codex goes from an agent that can write and review code to an agent that can do nearly anything developers and professionals can do on a computer." The model supports work across the software lifecycle — debugging, deploying, monitoring, writing PRDs, editing copy, user research, tests, and metrics — and extends to knowledge work beyond software, such as slide decks and spreadsheet analysis.
To test long-running agentic capabilities, OpenAI asked the model to build two games: version two of the racing game from the Codex app launch, and a diving game. Using generic follow-up prompts like "fix the bug" or "improve the game," GPT-5.3-Codex iterated on the games autonomously over millions of tokens. The company says the model also better understands intent on everyday websites than GPT-5.2-Codex, defaulting to more functional pages — for example, showing a yearly plan as a discounted monthly price and building an automatically transitioning testimonial carousel with three user quotes.
The Codex app now supports real-time steering. "Much like a colleague, you can steer and interact with GPT-5.3-Codex while it's working, without losing context," OpenAI writes. Users can enable steering in Settings > General > Follow-up behavior.
The model that helped build itself
The most notable claim in the announcement concerns self-improvement. OpenAI states that early versions of GPT-5.3-Codex were used to debug the model's own training, manage its own deployment, and diagnose test results — "our team was blown away by how much Codex was able to accelerate its own development." The company says many researchers and engineers describe their job today as fundamentally different from two months ago.
Concrete examples from the release: the research team used Codex to monitor and debug the training run; the engineering team used it to identify context rendering bugs and root-cause low cache hit rates; and the model is dynamically scaling GPU clusters during the launch to handle traffic surges. In one alpha-testing episode, a researcher asked the model to measure per-turn productivity, and GPT-5.3-Codex wrote regex classifiers, ran them over all session logs, and produced a report concluding that users were happier with fewer clarifying questions. A data scientist co-analyzed anomalous alpha results with the model, which summarized key insights over thousands of data points in under three minutes.
Cybersecurity guardrails
The High-capability cyber classification comes with new restrictions. OpenAI says it has "directly trained" the model to identify software vulnerabilities and, while it has no definitive evidence the model can automate cyber attacks end-to-end, it is deploying safety training, automated monitoring, trusted access for advanced capabilities, and enforcement pipelines including threat intelligence. Some requests flagged as elevated cyber risk will be automatically routed from GPT-5.3-Codex to GPT-5.2, with an appeals path through a new Trusted Access for Cyber pilot program.
OpenAI is also expanding the private beta of Aardvark, its security research agent, and partnering with open-source maintainers on free codebase scanning — including Next.js, where a researcher used Codex to find vulnerabilities disclosed last week (CVE-2025-59471 and CVE-2025-59472). Building on its $1M Cybersecurity Grant Program from 2023, the company is committing $10M in API credits for cyber defense work, especially for open-source software and critical infrastructure.
Availability
GPT-5.3-Codex is available on paid ChatGPT plans everywhere Codex runs: the app, CLI, IDE extension, and web. API access is "soon," with safety work underway. The model was co-designed for, trained with, and served on NVIDIA GB200 NVL72 systems.
OpenAI's stated trajectory is explicit: what started as the best coding agent is becoming "a more general collaborator on the computer," expanding both who can build software and what an agent can execute end to end.
Original: vercel.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI's GPT-5.1-Codex-Max Flags Coming Cybersecurity Threshold
- OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context
- OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
- OpenAI Takes Codex Coding Agent to General Availability
- OpenAI unveils GPT-5-Codex, a coding-tuned variant of GPT-5