Companies

OpenAI ships a model-native harness and native sandboxes for its Agents SDK

OpenAI's Agents SDK update adds a model-native harness, native sandbox execution across seven providers, and durable checkpointing — generally available now at standard API pricing.

The next evolution of the Agents SDK
The next evolution of the Agents SDKAI-generated
By Rebecca Stone5 min read

Updated

Why it matters

  • The updated Agents SDK is generally available to all customers via the API, with pricing based on standard token and tool-use rates.
  • Native sandbox execution supports Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel, plus a Manifest abstraction for AWS S3, Google Cloud Storage, Azure Blob Storage, and Cloudflare R2.
  • The new harness and sandbox capabilities launch first in Python; TypeScript support and additional features including code mode and subagents are planned for future releases.

OpenAI has released a major update to its Agents SDK, giving developers a model-native harness for agents that work across files, tools, and whole computer systems, plus native sandbox execution for running that work safely. The capabilities are generally available to all customers via the API and use standard API pricing, based on tokens and tool use.

The release lands at a moment when the gap between an impressive agent demo and a reliable production system remains wide. In its announcement, OpenAI framed the problem plainly: developers need more than the best models to build useful agents — they need systems that support how agents inspect files, run commands, write code, and keep working across many steps.

Why OpenAI thinks the current tooling falls short

OpenAI's diagnosis of the agent development stack reads like a critique of its own competitors. Model-agnostic frameworks are flexible, the company argues, but do not fully utilize frontier models' capabilities. Model-provider SDKs can sit closer to the model but often lack enough visibility into the harness. Managed agent APIs simplify deployment but constrain where agents run and how they access sensitive data.

The Agents SDK update is OpenAI's bid to occupy the space between those tradeoffs: infrastructure that is standardized and easy to get started with, but built specifically for OpenAI models. The pitch to developers is a harness that is "turnkey yet flexible" — one that adapts to their own stack, including their choices of tool use, memory, and sandbox environment, rather than forcing every product into a single mold.

A more capable harness for the agent loop

The core of the release is an upgraded harness for the agent loop — the cycle of reasoning, tool calls, and state management that defines how an autonomous system completes a task. The updated harness now includes configurable memory, sandbox-aware orchestration, and what OpenAI describes as Codex-like filesystem tools.

Crucially, the SDK standardizes integrations with primitives that are becoming common across frontier agent systems. OpenAI names them explicitly:

  • Tool use via MCP, the Model Context Protocol
  • Progressive disclosure via skills
  • Custom instructions via AGENTS.md
  • Code execution using the shell tool
  • File edits using the apply patch tool

OpenAI says the harness will continue to incorporate new agentic patterns and primitives over time, so developers spend less time on core infrastructure updates and more time on the domain-specific logic that makes their agents useful.

There is also a performance argument embedded in the design. The harness helps developers unlock more of a frontier model's capability by aligning execution with the way those models perform best, according to OpenAI. The company claims this keeps agents closer to the model's natural operating pattern, improving reliability and performance on complex tasks — particularly when work is long-running or coordinated across a diverse set of tools and systems.

That claim matters because it is the clearest statement yet of OpenAI's strategic positioning: the harness itself is a differentiator, not just the model. If execution scaffolding is tuned to how OpenAI's models behave, developers who build on it get better results — and a stronger reason to stay within OpenAI's ecosystem rather than abstracting across providers.

Native sandbox execution, with an open door to third parties

The second half of the release addresses where agents actually run. Many useful agents need a workspace where they can read and write files, install dependencies, run code, and use tools safely. Until now, developers typically assembled that execution layer themselves.

The updated Agents SDK supports sandbox execution natively. Developers can bring their own sandbox or use built-in support for seven providers: Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, and Vercel.

To make those environments portable across providers, the SDK introduces a Manifest abstraction for describing the agent's workspace. Developers can mount local files, define output directories, and bring in data from storage providers including AWS S3, Google Cloud Storage, Azure Blob Storage, and Cloudflare R2.

OpenAI frames the Manifest as a consistency mechanism: it gives developers a uniform way to shape the agent's environment from local prototype to production deployment. It also gives the model a predictable workspace — where to find inputs, where to write outputs, and how to keep work organized across a long-running task.

The inclusion of seven third-party sandbox providers is notable. Rather than locking execution to OpenAI infrastructure, the company is betting that standardized workspace descriptions — not proprietary compute — will anchor developers to its SDK.

Security by architectural separation

The most consequential design decision may be the separation of harness from compute. OpenAI's guidance is blunt: agent systems should be designed assuming prompt-injection and exfiltration attempts. Splitting the harness from the compute layer helps keep credentials out of environments where model-generated code executes.

That separation also enables durable execution. When the agent's state is externalized, losing a sandbox container does not mean losing the run. The SDK includes built-in snapshotting and rehydration, so it can restore the agent's state in a fresh container and continue from the last checkpoint if the original environment fails or expires.

Scale benefits follow the same architecture. Agent runs can use one sandbox or many, invoke sandboxes only when needed, route subagents to isolated environments, and parallelize work across containers for faster execution.

The security posture reflects a hardening consensus across the industry: as agents gain filesystem and shell access, the execution environment becomes the primary attack surface. OpenAI's answer is to treat credential isolation and state externalization as default SDK behavior rather than developer-implemented afterthoughts.

What comes next

The new harness and sandbox capabilities launch first in Python, with TypeScript support planned for a future release. OpenAI is also working to bring additional agent capabilities — including code mode and subagents — to both Python and TypeScript.

The company signaled broader ecosystem ambitions as well: support for more sandbox providers, more integrations, and more ways for developers to plug the SDK into the tools and systems they already use.

OpenAI says it will keep expanding the Agents SDK with the goal of making it easier to bring more capable agents into production with less custom infrastructure, while preserving the flexibility and control developers need to fit agents into their own environments. For teams currently stitching together frameworks, execution environments, and checkpointing logic, the release consolidates much of that stack into one SDK — and tightens the link between OpenAI's models and the infrastructure built to run them.

Original: developers.openai.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Ships AgentKit, Expanded Evals, and Reinforcement Fine-Tuning for Agents
  2. OpenAI turns the Responses API into a full agent runtime
  3. OpenAI Models Are Coming to Amazon Bedrock via a Stateful Agent Runtime
  4. OpenAI details codex-1: an o3 variant tuned for real coding work
  5. OpenAI details how it runs Codex safely in production

« Previous articleNext article »