Products & Tools

OpenAI Opens Up the Codex Harness via the App Server

OpenAI has detailed the Codex App Server, a bidirectional JSON-RPC protocol that lets any client embed the full Codex agent harness — IDEs, web, desktop and soon the CLI itself.

By James Calloway9 min read

Updated

Why it matters

  • All Codex surfaces — web app, CLI, IDE extension and the new macOS app — run on the same harness, exposed through the Codex App Server JSON-RPC API.
  • The App Server protocol is built on three primitives: items, turns and threads, with events like item/started, delta and completed.
  • Clients including JetBrains, Xcode and the Codex desktop app embed the harness; bindings exist in Go, Python, TypeScript, Swift and Kotlin.
  • OpenAI plans to refactor the Codex CLI TUI to speak the App Server protocol, enabling remote-machine sessions that survive laptop sleep.
  • The full source code is available in the open-source Codex CLI repo on GitHub.

OpenAI has pulled back the curtain on the Codex App Server, the bidirectional JSON-RPC protocol that now powers every Codex surface — the web app, the CLI, the IDE extension and the new Codex macOS app — and it is inviting external developers to build on it too.

In an engineering blog post titled "Unlocking the Codex harness: how we built the App Server," OpenAI explains that all of these products share a single underlying engine, the Codex harness: the agent loop and logic that drives every Codex experience. The App Server is the layer that exposes that harness to clients, and OpenAI now describes it as the first-class integration method it will maintain going forward.

The stakes are straightforward. As AI coding agents move from chat windows into IDEs, CI pipelines and desktop apps, whoever defines the stable protocol for embedding an agent wins the integration surface. OpenAI is positioning the App Server as that protocol for Codex — and publishing its source code in the open-source Codex CLI repository on GitHub.

Where did the App Server come from?

The App Server began as an internal convenience, not a platform bet. Codex CLI started life as a TUI, a terminal user interface. When OpenAI built the VS Code extension, engineers needed to drive the same agent loop from an IDE UI without re-implementing it.

That requirement went beyond simple request/response. The IDE needed to explore the workspace, stream progress as the agent reasoned, and emit diffs. OpenAI first experimented with exposing Codex as an MCP server, but the company says maintaining MCP semantics in a way that made sense for VS Code proved difficult.

Instead, the team introduced a JSON-RPC protocol that mirrored the TUI loop. That became the unofficial first version of the App Server. At the time, OpenAI did not expect other clients to depend on it, so it was not designed as a stable API.

Adoption changed the calculus. Internal teams and external partners wanted to embed the same harness in their own products. JetBrains and Xcode wanted an IDE-grade agent experience, while the Codex desktop app needed to orchestrate many Codex agents in parallel. Those demands pushed OpenAI to design a platform surface that its products and partner integrations could safely depend on over time — easy to integrate and backward compatible, so the protocol could evolve without breaking existing clients.

What is inside the Codex harness?

The harness is more than the core agent loop that orchestrates interaction between the user, the model and the tools. OpenAI breaks the full agent experience into three parts:

  • Thread lifecycle and persistence. A thread is a Codex conversation between a user and an agent. Codex creates, resumes, forks and archives threads, and persists the event history so clients can reconnect and render a consistent timeline.
  • Config and auth. Codex loads configuration, manages defaults and runs authentication flows like "Sign in with ChatGPT," including credential state.
  • Tool execution and extensions. Codex executes shell and file tools in a sandbox and wires up integrations like MCP servers and skills so they can participate in the agent loop under a consistent policy model.

All of this agent logic lives in a part of the Codex CLI codebase called "Codex core" — both a library where the agent code lives and a runtime that can be spun up to run the agent loop and manage the persistence of one Codex thread.

The App Server is the access layer: it is both the JSON-RPC protocol between client and server and a long-lived process that hosts the Codex core threads. An App Server process has four main components: the stdio reader, the Codex message processor, the thread manager and the core threads. The thread manager spins up one core session per thread, and the message processor communicates with each core session directly to submit client requests and receive updates.

One client request can result in many event updates, and those detailed events are what allow rich UIs to be built on top. The stdio reader and message processor act as a translation layer: they turn client JSON-RPC requests into Codex core operations, listen to Codex core's internal event stream, and transform those low-level events into a small set of stable, UI-ready JSON-RPC notifications.

The protocol is fully bidirectional. A typical thread has one client request and many server notifications. The server can also initiate requests when the agent needs input, such as an approval, and pause the turn until the client responds.

What are the conversation primitives?

Designing an API for an agent loop is tricky, OpenAI writes, because the user-agent interaction is not simple request/response. One user request can unfold into a structured sequence of actions the client must represent faithfully: the user's input, the agent's incremental progress, and artifacts produced along the way, such as diffs.

The protocol lands on three primitives with clear boundaries and lifecycles:

  • Item. The atomic unit of input/output in Codex. Items are typed — user message, agent message, tool execution, approval request, diff — and each has an explicit lifecycle: item/started when the item begins, optional item/*/delta events as content streams in for streaming item types, and item/completed when the item finalizes with its terminal payload. Clients can start rendering immediately on started, stream incremental updates on delta and finalize on completed.
  • Turn. One unit of agent work initiated by user input. It begins when the client submits an input — for example, "run tests and summarize failures" — and ends when the agent finishes producing outputs. A turn contains a sequence of items representing intermediate steps and outputs.
  • Thread. The durable container for an ongoing Codex session. It contains multiple turns, can be created, resumed, forked and archived, and its history is persisted so clients can reconnect and render a consistent timeline.

Every conversation begins with an initialize handshake: the client must send a single initialize request before any other method, and the server responds by advertising capabilities and agreeing on protocol versioning, feature flags and defaults. When the client makes a new request, it creates a thread and then a turn, and the server sends back notifications for progress (thread/started and turn/started) plus items such as the user message.

Tool calls also come back to the client as items. The server may ask for client approval before running an action — OpenAI shows a VS Code permission prompt reading "Do you want to allow me to run pnpm test for this workspace?" — and the approval pauses the turn until the client replies with either "allow" or "deny." The turn ends with turn/completed, with agent message deltas streaming until the message is finalized with item/completed.

How do clients integrate?

Across all client surfaces, the transport is JSON-RPC over stdio (JSONL). Codex surfaces and partner integrations have implemented App Server clients in Go, Python, TypeScript, Swift and Kotlin. For TypeScript, developers can generate definitions directly from the Rust protocol; for other languages, they can generate a JSON Schema bundle and feed it into their preferred code generator.

OpenAI describes three integration patterns:

  • Local apps and IDEs. Local clients bundle or fetch a platform-specific App Server binary, launch it as a long-running child process and keep a bidirectional stdio channel open. OpenAI's VS Code extension and Desktop App pin the shipped artifact to a tested version so the client always runs the exact bits the company validated. Partners like Xcode decouple release cycles differently: the client stays stable and points to a newer App Server binary when needed, adopting server-side improvements — better auto-compaction in Codex core, newly supported config keys — without waiting for a client release. The JSON-RPC surface is backward compatible, so older clients can talk to newer servers safely.
  • Codex Web. The web product runs the harness in a container environment. A worker provisions a container with the checked-out workspace, launches the App Server binary inside it and maintains a long-lived JSON-RPC-over-stdio channel. The browser app talks to the Codex backend over HTTP and SSE, which streams task events produced by the worker. Because web sessions are ephemeral, the web app cannot be the source of truth for long-running tasks; state and progress stay on the server, so work continues even if the tab disappears. A new session can reconnect, pick up where it left off and catch up without rebuilding client state.
  • TUI/Codex CLI. Historically the TUI was a native client that ran in the same process as the agent loop and talked directly to Rust core types. OpenAI now plans to refactor the TUI to use the App Server so it behaves like any other client: launch a child process, speak JSON-RPC over stdio, render the same streaming events and approvals. That unlocks workflows where the TUI connects to a Codex server on a remote machine, keeping the agent close to compute and continuing work even if the laptop sleeps or disconnects, while still delivering live updates and controls locally.

Which protocol should developers choose?

OpenAI recommends the App Server by default, but the blog post maps the alternatives:

  • Codex as an MCP server. Run codex mcp-server and connect from any MCP client that supports stdio servers, such as the OpenAI Agents SDK. A good fit for existing MCP-based workflows that want Codex as a callable tool. The downside: you only get what MCP exposes, so Codex-specific interactions that rely on richer session semantics, such as diff updates, may not map cleanly.
  • Cross-provider agent harness protocols. Portable interfaces that target multiple model providers and runtimes. Useful for coordinating multiple agents under one abstraction, but these protocols often converge on the common subset of capabilities, which can make richer interactions harder to represent. OpenAI notes this space is evolving quickly and expects more common standards to emerge, citing skills as a good example.
  • Codex App Server. The full harness exposed as a stable, UI-friendly event stream, including Sign in with ChatGPT, model discovery and configuration management. The main cost is integration work: building the client-side JSON-RPC binding in your language. In practice, OpenAI says Codex itself can do much of the heavy lifting if you feed it the JSON schema and documentation, and many teams the company worked with reached a working integration quickly using Codex.

Two lighter options round out the lineup: a scriptable CLI mode for one-off tasks and CI runs that exits with a clear success or failure signal, and a TypeScript library for programmatically controlling local Codex agents from within an application. The library shipped earlier than the App Server and currently supports fewer languages and a smaller surface area; OpenAI says it may add additional SDKs wrapping the App Server protocol if there is developer interest.

The move signals where OpenAI thinks agent infrastructure is heading: stable, versioned protocols that third parties can build products on, rather than one-off integrations. With the App Server documented, backward compatible and open source, the question for developers embedding AI coding agents is no longer whether they can hook into Codex — it is whether OpenAI's primitives become the default way the industry wires agents into its tools.

Original: chatgpt.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

215 articles

Related articles

  1. OpenAI Ships Codex Desktop App as Agent Command Center
  2. OpenAI Turns Codex Cloud Environments Persistent and Reusable Across Devices
  3. OpenAI says Codex runs 'daily' across its engineering teams
  4. OpenAI Upgrades Codex: Faster, More Reliable, More Autonomous
  5. OpenAI to Acquire Ona, Pushing Codex Toward Persistent Cloud Agents

« Previous articleNext article »