Enterprise & Work

OpenAI Shipped a Million Lines of Code With Zero Human-Written Lines

OpenAI engineers shipped a product with ~1M lines of code, all written by Codex, in ~1/10th the hand-coded time. Humans designed environments; agents executed.

Harness engineering: leveraging Codex in an agent-first world
Harness engineering: leveraging Codex in an agent-first worldstriatic / Openverse
By Rebecca Stone6 min read

Updated

Why it matters

  • Every line of the product's code — including tests, CI config, docs, and tooling — was written by Codex, with zero manually-written code over five months.
  • Three engineers merged ~1,500 PRs (3.5 PRs/engineer/day), producing on the order of a million lines of code; throughput increased as the team grew to seven.
  • The repository reached end-to-end autonomy: from a single prompt, Codex can reproduce a bug, record video evidence, fix, validate, open a PR, respond to reviews, and merge — but the team warns this doesn't yet generalize without similar investment.

A team at OpenAI has built and shipped an internal software product over the past five months in which every single line of code — application logic, tests, CI configuration, documentation, observability, and internal tooling — was written by Codex, not humans. The team estimates the approach compressed roughly a million lines of code into about one-tenth of the time hand-writing it would have required.

The first commit to an empty repository landed in late August 2025, generated by Codex CLI running GPT-5 and guided by a small set of existing templates. Even the initial AGENTS.md file, which directs agents on how to work in the repository, was itself written by Codex. "Humans steer. Agents execute," the team writes.

Three engineers drove Codex through roughly 1,500 merged pull requests, an average throughput of 3.5 PRs per engineer per day. Throughput increased as the team grew to seven engineers. The product is not a toy: it has internal daily users, external alpha testers, and has been used by hundreds of users internally. It ships, deploys, breaks, and gets fixed.

The engineer's job changes

Early progress was slower than expected — not because Codex was incapable, the team says, but because the environment was underspecified. The agent lacked the tools, abstractions, and internal structure required to make progress toward high-level goals. "The primary job of our engineering team became enabling the agents to do useful work."

When something failed, the fix was almost never "try harder." Engineers instead asked: "what capability is missing, and how do we make it both legible and enforceable for the agent?" Humans interact with the system almost entirely through prompts, instructing Codex to review its own changes, request additional agent reviews locally and in the cloud, and iterate until all agent reviewers are satisfied — a loop the team links to the "Ralph Wiggum Loop" pattern. Human review still happens, but the team has pushed almost all of it toward agent-to-agent review.

Making the application legible to the agent

As code throughput rose, human QA capacity became the bottleneck. The team's response was to make the application itself directly legible to Codex. They made the app bootable per git worktree, so Codex could launch one instance per change, and wired the Chrome DevTools Protocol into the agent runtime, creating skills for DOM snapshots, screenshots, and navigation. Codex can now reproduce bugs, validate fixes, and reason about UI behavior directly.

Observability got the same treatment. Logs, metrics, and traces are exposed to Codex via an ephemeral local observability stack, queryable with LogQL and PromQL. That makes prompts like "ensure service startup completes in under 800ms" tractable. Single Codex runs regularly work on a single task for upwards of six hours — often while the humans sleep.

A map, not a manual

One of the earliest lessons concerned context management: "give Codex a map, not a 1,000-page instruction manual." The team tried the "one big AGENTS.md" approach and says it failed in predictable ways: context is a scarce resource that a giant file crowds out; too much guidance becomes non-guidance; monolithic manuals "rot instantly"; and a single blob is hard to verify mechanically.

Instead, AGENTS.md — roughly 100 lines — now serves as a table of contents, with the knowledge base living in a structured docs/ directory treated as the system of record. Design documentation is catalogued and indexed with verification status. A quality document grades each product domain and architectural layer, tracking gaps over time. Plans are first-class artifacts: complex work is captured in execution plans with progress and decision logs checked into the repository. Dedicated linters and CI jobs validate that the knowledge base stays current, and a recurring "doc-gardening" agent opens fix-up pull requests for stale documentation.

Boring technology wins

Because the repository is entirely agent-generated, it is optimized first for Codex's legibility. "From the agent's point of view, anything it can't access in-context while running effectively doesn't exist," the team writes. Knowledge that lives in Google Docs, chat threads, or people's heads is invisible to the system — so the team pushes more and more context into the repo over time.

This framing favored "boring" technologies, which tend to be easier for agents to model due to composability, API stability, and representation in the training set. In some cases it was cheaper to have the agent reimplement functionality than to work around opaque upstream behavior. Rather than pulling in a generic p-limit-style package, the team implemented its own map-with-concurrency helper — tightly integrated with its OpenTelemetry instrumentation, with 100% test coverage.

Enforcing invariants, not micromanaging

Documentation alone doesn't keep a fully agent-generated codebase coherent. The team enforces strict architectural boundaries: each business domain is divided into a fixed set of layers with strictly validated dependency directions (Types → Config → Repo → Service → Runtime → UI), enforced mechanically via custom linters and structural tests — themselves Codex-generated. "This is the kind of architecture you usually postpone until you have hundreds of engineers. With coding agents, it's an early prerequisite."

Custom lint error messages are written to inject remediation instructions directly into agent context. Human taste is fed back continuously: review comments, refactoring PRs, and user-facing bugs are captured as documentation updates or encoded directly into tooling. "When documentation falls short, we promote the rule into code."

Garbage collection for AI slop

Full autonomy introduced a novel problem: Codex replicates patterns that already exist in the repository, even suboptimal ones, leading to drift. The team initially spent every Friday — 20% of the week — manually cleaning up "AI slop." That didn't scale. Instead, they encoded "golden principles" into the repository and run recurring background Codex tasks that scan for deviations, update quality grades, and open targeted refactoring pull requests, most reviewable in under a minute and automerged.

"This functions like garbage collection," the team writes. "Technical debt is like a high-interest loan: it's almost always better to pay it down continuously in small increments."

End-to-end autonomy, with caveats

The repository has crossed a threshold where Codex can drive a new feature end-to-end from a single prompt: validate the codebase, reproduce a bug, record a video of the failure, implement a fix, validate it by driving the app, record a resolution video, open a PR, respond to feedback, remediate build failures, escalate to humans only when judgment is required, and merge. The team cautions that this behavior "depends heavily on the specific structure and tooling of this repository and should not be assumed to generalize without similar investment — at least, not yet."

Open questions remain. The team doesn't know how architectural coherence evolves over years in a fully agent-generated system, where human judgment adds the most leverage, or how the system will change as models improve. What has become clear, they say, is that building software still demands discipline — but the discipline now shows up in the scaffolding rather than the code. As agents take on larger portions of the software lifecycle, the hardest engineering challenges will center on "designing environments, feedback loops, and control systems" that let agents build and maintain complex, reliable software at scale.

Original: ghuntley.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI to Acquire Ona, Pushing Codex Toward Persistent Cloud Agents
  2. OpenAI Takes Codex Coding Agent to General Availability
  3. OpenAI Open-Sources Symphony, a Spec That Turns Linear Into an Agent Orchestrator
  4. OpenAI Upgrades Codex: Faster, More Reliable, More Autonomous
  5. OpenAI's Codex Hits 4 Million Weekly Developers, Launches Enterprise Push

« Previous articleNext article »