OpenAI built an internal data agent on GPT-5.2 and Codex
OpenAI detailed an internal GPT-5.2 data agent that serves 3,500 users across 600 petabytes, built from Codex, Evals, and embeddings — tools the company also sells to developers.

Updated
Why it matters
- OpenAI's internal data agent runs on GPT-5.2 and serves 3,500+ users working across 600 petabytes of data and 70,000 datasets.
- The agent is built on six layers of context: table usage, human annotations, Codex enrichment, institutional knowledge, memory, and runtime context, retrieved via RAG with OpenAI's embeddings API.
- Quality is protected by continuous Evals API testing against manually authored 'golden' SQL queries, and all data access is strictly pass-through to existing user permissions.
OpenAI has built an in-house AI data agent, powered by its GPT‑5.2 flagship model, that lets employees across Engineering, Data Science, Go-To-Market, Finance, and Research go "from question to insight in minutes, not days." The company detailed the internal-only tool — not an external product — in a post describing how it was constructed from the same building blocks it sells to developers: Codex, GPT‑5, the Evals API, and the Embeddings API.
The stakes are scale. OpenAI's data platform serves more than 3,500 internal users working across Engineering, Product, and Research, spanning over 600 petabytes of data across 70,000 datasets. At that size, the company writes, simply finding the right table can be one of the most time-consuming parts of doing analysis.
One internal user described the problem bluntly: "We have a lot of tables that are fairly similar, and I spend tons of time trying to figure out how they're different and which to use. Some include logged-out users, some don't. Some have overlapping fields; it's hard to tell what is what."
Even with correct tables selected, common SQL failure modes — many-to-many joins, filter pushdown errors, unhandled nulls — can silently invalidate results. OpenAI cites one internal SQL statement that runs more than 180 lines long, making it hard to verify that joins and columns are correct.
How the agent works
The agent runs on GPT‑5.2 and is available wherever employees already work: as a Slack agent, through a web interface, inside IDEs, in the Codex CLI via MCP, and directly in OpenAI's internal ChatGPT app through an MCP connector.
OpenAI says the agent handles analysis end-to-end — from understanding the question to exploring the data, running queries, and synthesizing findings. It gave a test-data example prompt: "For NYC taxi trips, which pickup-to-dropoff ZIP pairs are the most unreliable, with the largest gap between typical and worst-case travel times, and when does that variability occur?"
Rather than following a fixed script, the agent evaluates its own progress. If an intermediate result looks wrong — a zero-row result from an incorrect join, for instance — it investigates, adjusts its approach, and retries, carrying learnings forward between steps. OpenAI describes this as a "closed-loop, self-learning process" that shifts iteration from the user into the agent itself. The agent covers the full analytics workflow: discovering data, running SQL, and publishing notebooks and reports. It can also web search for external information.
Six layers of context
High-quality answers depend on rich context, OpenAI writes. Without it, even strong models can misestimate user counts or misinterpret internal terminology. The agent is built around six context layers.
Layer 1: Table usage. The agent uses schema metadata — column names and data types — to inform SQL writing, plus table lineage to understand upstream and downstream relationships. Ingested historical queries teach it which tables are typically joined together.
Layer 2: Human annotations. Curated descriptions of tables and columns from domain experts capture intent, semantics, business meaning, and known caveats that can't be inferred from schemas or past queries alone.
Layer 3: Codex enrichment. By deriving a code-level definition of a table with Codex, the agent builds a deeper understanding of what the data actually contains — uniqueness of values, update frequency, scope, and whether a table includes only first-party ChatGPT traffic. This context refreshes automatically, without manual maintenance.
Layer 4: Institutional knowledge. The agent accesses Slack, Google Docs, and Notion, which capture launches, reliability incidents, internal codenames, and canonical metric definitions. Documents are ingested, embedded, and stored with metadata and permissions; a retrieval service handles access control and caching at runtime.
Layer 5: Memory. When the agent receives corrections or discovers nuances, it saves those learnings for next time. OpenAI gives a concrete example: in one case, the agent didn't know how to filter for a particular analytics experiment that relied on matching a specific string defined in an experiment gate. Memory let it filter correctly instead of fuzzily string-matching. Memories are scoped at global and personal levels, can be manually created and edited, and the agent prompts users to save learnings it finds.
Layer 6: Runtime context. When no prior context exists or information is stale, the agent issues live queries to the data warehouse to inspect tables directly, and can talk to other data platform systems — metadata service, Airflow, Spark — for broader context.
A daily offline pipeline aggregates table usage, annotations, and Codex-derived enrichment into a single normalized representation. That context is converted into embeddings using OpenAI's embeddings API and stored. At query time, the agent pulls only relevant embedded context via retrieval-augmented generation rather than scanning raw metadata or logs, keeping runtime latency low even across tens of thousands of tables.
Built like a teammate
The agent is conversational and always-on, carrying complete context across turns. Users can interrupt mid-analysis and redirect it. When instructions are unclear, it asks clarifying questions; if no response comes, it applies sensible defaults — a growth question with no date range might assume the last seven or 30 days.
After rollout, OpenAI observed users repeatedly running the same routine analyses. The agent's workflows now package recurring analyses — weekly business reports, table validations — into reusable instruction sets, encoding context and best practices once for consistent results across users.
Guarding quality with evals
Building an always-on, evolving agent means quality can drift as easily as it can improve. OpenAI uses its Evals API to measure and protect response quality. Each eval pairs a question targeting an important metric or analytical pattern with a manually authored "golden" SQL query that produces the expected result. The natural-language question goes to the query-generation endpoint, the generated SQL executes, and the output is compared against the expected result.
Evaluation does not rely on string matching. Generated SQL can differ syntactically while remaining correct. OpenAI compares both the SQL and resulting data, feeding the signals into the Evals grader, which produces a final score plus an explanation. These evals run continuously, like unit tests, catching regressions as "canaries in production."
Security model
The agent plugs into OpenAI's existing security and access-control model, operating purely as an interface layer. All access is strictly pass-through: users can only query tables they already have permission to access, and when access is missing, the agent flags it or falls back to authorized datasets. It exposes its reasoning by summarizing assumptions and execution steps alongside each answer, and links directly to underlying query results for inspection.
Three lessons learned
OpenAI says building the agent from scratch surfaced practical lessons about agent reliability at scale.
Less is more. Exposing the full tool set early caused problems with overlapping functionality. Redundancy that is manageable for humans is confusing to agents, so OpenAI restricted and consolidated tool calls.
Guide the goal, not the path. Highly prescriptive prompting degraded results. Rigid instructions often pushed the agent down incorrect paths; higher-level guidance plus GPT‑5's reasoning produced better outcomes.
Meaning lives in code. Schemas and query history describe a table's shape and usage, but its true meaning lives in the code that produces it — pipeline logic captures assumptions, freshness guarantees, and business intent that never surface in SQL or metadata. By crawling the codebase with Codex, the agent answers "what's in here" and "when can I use it" far more accurately than from warehouse signals alone.
What's next
OpenAI says it is improving the agent's handling of ambiguous questions, strengthening validations, and integrating it more deeply into workflows, with the goal that it "blend naturally into how people already work, instead of functioning like a separate tool." While the tooling will benefit from improvements in agent reasoning, validation, and self-correction, the team's mission stays fixed: seamlessly deliver fast, trustworthy data analysis across OpenAI's data ecosystem.
The post doubles as a demonstration for OpenAI's commercial stack. Every component behind the agent — Codex, GPT‑5, the Evals API, the Embeddings API, MCP connectors — is available to developers today, and the writeup reads as an implicit blueprint for enterprises weighing whether to build the same class of tool on top of it.
Original: platform.openai.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- OpenAI Launches ChatGPT for Financial Services With Built-In Data
- OpenAI Puts ChatGPT Inside Excel With GPT-5.4 Built for Finance
- OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
- OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work