Models

OpenAI's GPT-5.1-Codex-Max runs coding tasks for 24 hours straight

OpenAI's GPT-5.1-Codex-Max uses compaction to work across millions of tokens, ran tasks for 24+ hours internally, and now replaces GPT-5.1-Codex as Codex's default model.

Building more with GPT-5.1-Codex-Max
Building more with GPT-5.1-Codex-Maxjon_a_ross / Openverse
By Sophie Lindqvist4 min read

Updated

Why it matters

  • GPT-5.1-Codex-Max is OpenAI's first model natively trained to work across multiple context windows via compaction, handling millions of tokens in a single task; internal evaluations observed it working for more than 24 hours.
  • It scores 77.9% on SWE-bench Verified (n=500) versus 73.7% for GPT-5.1-Codex, and at 'medium' reasoning effort beats its predecessor while using 30% fewer thinking tokens.
  • The model replaces GPT-5.1-Codex as the default in Codex today for ChatGPT Plus, Pro, Business, Edu, and Enterprise plans; API access is coming soon.

OpenAI has launched GPT-5.1-Codex-Max, a frontier agentic coding model that it says worked on single tasks for more than 24 hours in internal evaluations — and it replaces GPT-5.1-Codex as the default model across Codex starting today.

The model is OpenAI's first trained to operate across multiple context windows through a process the company calls compaction. When a session approaches the context-window limit, the model prunes its history while preserving the most important context, resets to a fresh window, and repeats the process until the task is done. According to OpenAI, this lets the model coherently work over millions of tokens in a single task, unlocking project-scale refactors, deep debugging sessions, and multi-hour agent loops that previously would have failed on context limits.

OpenAI frames that persistence as the point. "The ability to sustain coherent work over long horizons is a foundational capability on the path toward more general, reliable AI systems," the company writes. In internal evaluations, GPT-5.1-Codex-Max "will persistently iterate on its implementation, fix test failures, and ultimately deliver a successful result" — demonstrated in a video where the model independently refactors the Codex CLI open-source repository, compacting its session as it goes.

Benchmark gains and lower token bills

The model outperforms its predecessor on frontier coding benchmarks. On SWE-bench Verified (n=500), GPT-5.1-Codex-Max at the new Extra High ('xhigh') reasoning effort scores 77.9%, versus 73.7% for GPT-5.1-Codex at high effort. On SWE-Lancer IC SWE, the gap is wider: 79.9% versus 66.3%. Terminal-Bench 2.0 improves from 52.8% to 58.1%.

Efficiency moved alongside accuracy. OpenAI reports that GPT-5.1-Codex-Max at 'medium' reasoning effort beats GPT-5.1-Codex at the same setting while using 30% fewer thinking tokens. The company expects those savings to reach developers' invoices: it cites frontend generation as an example, where the new model produces designs with "similar functionality and aesthetics" at "much lower cost." OpenAI still recommends medium effort as the daily driver for most tasks; xhigh exists for non-latency-sensitive work.

Training also changed. OpenAI says the model was trained on real-world software engineering tasks — PR creation, code review, frontend coding, and Q&A — and is the first model it has trained to operate in Windows environments. Training now also includes tasks designed to make the model a better collaborator in the Codex CLI.

Cybersecurity posture and guardrails

The capabilities come with security caveats. OpenAI states that GPT-5.1-Codex-Max "does not reach High capability on Cybersecurity under our Preparedness Framework," but calls it "the most capable cybersecurity model we've deployed to date." Because agentic cyber capabilities are evolving quickly, the company says it is preparing mitigations for High capability and working to route defensive benefits to security teams through programs like Aardvark.

Since the GPT-5-Codex launch, OpenAI says it has run dedicated cybersecurity-specific monitoring. It has "not observed a meaningful increase in scaled abuse," but its teams have already disrupted cyber operations attempting to misuse its models, with suspicious activity routed through policy monitoring systems. Full first- and third-party evaluation results appear in the model's system card.

Codex itself runs in a secure sandbox by default: file writes stay within the workspace, and network access stays off unless a developer enables it. OpenAI recommends keeping that restricted mode, since web access can introduce prompt-injection risks from untrusted content. On deployment, the company is explicit about limits: Codex logs terminal output and cites tool calls and test results, and its code reviews reduce the risk of shipping model- or human-produced bugs, but "Codex should be treated as an additional reviewer and not a replacement for human reviews."

Availability and internal traction

GPT-5.1-Codex-Max is available now in Codex across the CLI, IDE extension, cloud, and code review for ChatGPT Plus, Pro, Business, Edu, and Enterprise plans. API access for Codex CLI users is "coming soon." OpenAI cautions that unlike GPT-5.1, a general-purpose model, GPT-5.1-Codex-Max and the Codex family should be used only for agentic coding tasks in Codex or Codex-like environments.

The internal adoption numbers explain why OpenAI is defaulting everyone to the new model: 95% of OpenAI engineers use Codex weekly, and those engineers ship roughly 70% more pull requests since adopting the tool. With compaction pushing agents toward day-long autonomous work, the practical question for teams shifts from whether the model can finish a task to how they review what it produced before it reaches production.

Original: developers.openai.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

114 articles

Related articles

  1. OpenAI's GPT-5.1-Codex-Max Flags Coming Cybersecurity Threshold
  2. OpenAI unveils GPT-5-Codex, a coding-tuned variant of GPT-5
  3. OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
  4. OpenAI's WebSocket Overhaul Makes Agents 40% Faster
  5. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price

« Previous articleNext article »