Enterprise & Work

Datadog found OpenAI's Codex could have prevented 22% of past incidents

Datadog replayed Codex against pull requests behind real incidents and found the agent would have prevented roughly 22% of them. Over 1,000 engineers now use it for code review.

Datadog uses Codex for system-level code review
Datadog uses Codex for system-level code reviewElogia Marketing4eCommerce / Openverse
By Rebecca Stone4 min read

Updated

Why it matters

  • Datadog engineers confirmed Codex feedback would have made a difference in more than 10 historical incidents, roughly 22% of those examined — more than any other tool evaluated.
  • More than 1,000 Datadog engineers now use OpenAI's Codex regularly for code review.
  • Brad Carter, who leads Datadog's AI DevX team, says 'preventing incidents is far more compelling at our scale' than time savings.

Datadog's engineers confirmed that feedback from OpenAI's coding agent Codex would have made a difference in more than 10 historical incidents — roughly 22% of the outages the company replayed the tool against, more than any other tool it evaluated. The result comes from an incident replay harness the observability company built to test whether AI-assisted code review could do more than flag style issues.

The stakes are considerable. Datadog runs one of the world's most widely used observability platforms, and customers depend on it to surface problems when their own systems break. That makes code review a high-stakes moment for Datadog's engineering teams: it is not just about catching mistakes, but about understanding how changes ripple through interconnected distributed systems — an area where traditional static analysis and rule-based tools often fall short.

"Time savings are real and important," says Brad Carter, who leads Datadog's AI Development Experience (AI DevX) team. "But preventing incidents is far more compelling at our scale."

Testing against real incidents, not hypotheticals

Effective code review at Datadog traditionally relied on senior engineers — the people who know the codebase, its history, and its architectural tradeoffs well enough to spot systemic risk. That kind of deep context is hard to scale. Early AI code review tools did not solve the problem either; Datadog engineers found many behaved like advanced linters, flagging surface-level issues while missing broader system nuances, and ignored the noisy or shallow suggestions.

To measure Codex against that baseline, the AI DevX team reconstructed pull requests that had contributed to historical incidents, ran Codex against each one as if it were part of the original review, then asked the engineers who owned those incidents whether the feedback would have changed the outcome. Because those pull requests had already passed human review, the replay showed Codex surfacing risks reviewers had not seen at the time — complementing human judgment rather than replacing it.

Datadog began its pilot by integrating Codex into live development workflows. In one of the company's largest and most heavily used repositories, every pull request was automatically reviewed by the agent. Engineers reacted to comments with thumbs up or down, and many noted the feedback was worth reading — unlike suggestions from previous tools.

Feedback beyond the diff

Datadog's analysis showed Codex consistently flagged issues that are not obvious from the immediate diff alone and cannot be caught by deterministic rules. Engineers described the comments as more than "bot noise": the agent pointed out interactions with modules not touched in the diff, identified missing test coverage in areas of cross-service coupling, and highlighted API contract changes that carried downstream risk.

"For me, a Codex comment feels like the smartest engineer I've worked with and who has infinite time to find bugs," says Carter. "It sees connections my brain doesn't hold all at once."

Unlike static analysis tools, Codex compares the intent of the pull request with the submitted code changes, reasoning over the entire codebase and its dependencies, and executes code and tests to validate behavior. "It was the first one that actually seemed to consider the diff in the larger context of the program," says Carter. "That was novel and eye-opening."

That difference changed how engineers engaged with AI review. "I started treating Codex comments like real code review feedback," says Ted Wexler, Senior Software Engineer at Datadog. "Not something I'd skim or ignore, but something worth paying attention to."

From detection to design

Following the evaluation, Datadog deployed Codex more broadly. Today more than 1,000 engineers use it regularly, with feedback surfacing organically — engineers post to Slack about useful insights and moments where Codex helped them think differently about a problem.

The broader shift is in how Datadog defines code review itself: no longer a checkpoint for catching errors or optimizing cycle time, but a core reliability partner. Carter frames the new role as surfacing risk beyond what individual reviewers can hold in context, highlighting cross-module and cross-service interactions, and letting human reviewers focus on architecture and design.

"Codex changed my mind for what code review should be," says Carter. "It's not about replicating our best human reviewers. It's about finding critical flaws and edge cases that humans struggle to see when reviewing changes in isolation."

The deployment aligns with how Datadog's leaders set engineering priorities, where reliability and trust matter as much as — if not more than — velocity. As Carter puts it: "We are the platform companies rely on when everything else is breaking. Preventing incidents strengthens the trust our customers place in us."

For a company whose product is other companies' reliability, the 22% figure now serves as the benchmark any future review tool at Datadog will have to beat.

Original: datadoghq.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Says Internal AI Monitor Caught Every Employee-Reported Misuse Case
  2. OpenAI's Codex Security Enters Research Preview
  3. OpenAI Launches GPT-5.2-Codex, Its Most Advanced Coding Model
  4. OpenAI and Paradigm Launch EVMbench for Smart Contract Security
  5. OpenAI Tells Evaluators: The Harness Is Part of the Result

« Previous articleNext article »