OpenAI Tells Business Leaders: Write Evals, Not Wish Lists
OpenAI says one million businesses use AI but many see no ROI. Its fix: contextual evals — specify, measure, improve — built by cross-functional teams, not engineers alone.

Updated
Why it matters
- Over one million businesses worldwide use AI, according to OpenAI, but some struggle to achieve expected results.
- OpenAI's framework has three stages — specify, measure, improve — and recommends reviewing 50 to 100 early outputs to build an error taxonomy.
- OpenAI says contextual evals create a hard-to-copy, context-specific dataset and that "management skills are AI skills."
OpenAI says more than one million businesses worldwide now use AI, and it has a message for the ones not getting results: write evals. In a new guide published on its site, the company argues that organizations struggling to see ROI from AI are failing at measurement — and it lays out a framework for fixing that.
Evals, OpenAI explains, are "methods to measure and improve the ability of an AI system to meet expectations." The company frames them as the AI-era equivalent of product requirement documents: they turn fuzzy goals into explicit, testable criteria. Used strategically, the guide claims, evals make products and internal tools more reliable at scale, reduce high-severity errors, protect against downside risk, and create "a measurable path to higher ROI."
OpenAI practices what it preaches internally. Its researchers run rigorous frontier evals — published at evals.openai.com — to benchmark model performance across domains. But frontier evals have a limit: they "cannot reveal all the nuances required to ensure the model will perform on a specific workflow in a specific business setting." So OpenAI's internal teams have built dozens of contextual evals tailored to specific products and workflows, and the company says business leaders should learn to do the same.
The stakes are straightforward. As models commoditize and expertise spreads, differentiation shifts to execution inside a company's own context. OpenAI puts it bluntly: "In a world where information is freely available across the world and expertise is democratized, your advantage hinges on how well your systems can execute inside your context."
Specify, measure, improve
The framework OpenAI presents has three stages.
Specify. A small, cross-functional team — mixing technical and domain expertise — writes down the system's purpose in plain terms, for example: "Convert qualified inbound emails into scheduled demos while staying on brand." The team maps the workflow end-to-end, defines success and failure at each decision point, and builds a "golden set" of example inputs paired with desired outputs. OpenAI advises against trying to solve everything at once: reviewing 50 to 100 outputs from an early prototype surfaces how and where the system fails, producing a taxonomy of errors and their frequencies. The company stresses this is not a purely technical exercise — technical teams "should not be asked in isolation to judge what best serves customers."
Measure. Testing should happen in a dedicated environment that mirrors real-world conditions, not a prompt playground. Use real-world examples and invent edge cases that are rare but costly if mishandled. Rubrics help, but OpenAI warns against over-emphasizing superficial items. Some evals can scale with an LLM grader — an AI model that grades outputs the way an expert would — but a human must stay in the loop, auditing the grader and reviewing system logs directly. Measurement does not stop at launch: end-user signals, external or internal, should feed back into the eval.
Improve. Fixing what the eval uncovers can mean refining prompts, adjusting data access, or updating the eval itself. OpenAI recommends building a "data flywheel": log inputs, outputs, and outcomes; route ambiguous or costly cases to expert review on a schedule; fold those judgments back into prompts, tools, and models. The payoff, the company argues, is "a large, differentiated, context-specific dataset that is hard to copy."
Caveats and context
OpenAI is explicit that contextual evals remain immature. Definitive processes "have yet to emerge," the guide states, and evals must be continuously maintained and stress-tested as models, data, and business goals evolve. For external products, evals complement — not replace — A/B tests and product experimentation.
The company positions evals as the successor to OKRs and KPIs, the measurement frameworks of the big-data era. Its closing argument is managerial as much as technical: "management skills are AI skills." Leaders working with probabilistic systems must decide when precision is essential, when flexibility is acceptable, and how to balance velocity against reliability.
For teams building on its API, OpenAI points to its Platform Docs as a starting point. The guide ends with a line that doubles as its thesis: "Don't hope for 'great.' Specify it, measure it, and improve toward it." Expect more guidance from OpenAI as contextual eval practices mature — the company says it will share emerging frameworks as they develop.
Original: evals.openai.com
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI Launches GDPval, a Benchmark Built From Real Jobs
- OpenAI Tells Evaluators: The Harness Is Part of the Result
- OpenAI Puts $150 Million Into New Global Partner Network
- OpenAI Takes Equity Stake in Thrive Holdings to Push AI Into Enterprise Operations
- OpenAI Launches Frontier, an Enterprise Platform for AI Agents