OpenAI Launches GDPval, a Benchmark Built From Real Jobs
OpenAI's new GDPval benchmark tests models on 1,320 real professional tasks. Claude Opus 4.1 led, and models finished tasks ~100x faster and cheaper than experts.

Updated
Why it matters
- GDPval spans 44 occupations across 9 industries and includes 1,320 tasks, with 220 in an open-sourced gold set, vetted by professionals averaging 14+ years of experience.
- In blind expert grading, Claude Opus 4.1 was the best performing model, rated as good as or better than humans on just under half of the 220 gold-set tasks; performance more than tripled from GPT-4o to GPT-5.
- Frontier models completed GDPval tasks roughly 100x faster and 100x cheaper than industry experts, though those figures exclude human oversight, iteration, and integration costs.
OpenAI has released GDPval, a new evaluation that measures how well AI models perform on economically valuable, real-world knowledge work — and the first results show frontier models already matching industry professionals on nearly half the tasks, at roughly 100 times the speed and 100 times lower cost.
The benchmark, announced Thursday, spans 44 occupations drawn from the 9 industries that each contribute more than 5% of U.S. GDP, according to data from the Federal Reserve Bank of St. Louis. The full set contains 1,320 specialized tasks, with 220 tasks in an open-sourced "gold" set. Every task was written and vetted by experienced professionals who average more than 14 years in their fields.
The stakes are straightforward. Debate about AI's impact on the labor market runs mostly on speculation. OpenAI wants to replace that with measurement. "Evaluations like GDPval help ground conversations about future AI improvements in evidence rather than guesswork, and can help us track model improvement over time," the company wrote in its announcement.
A benchmark based on deliverables, not exam questions
GDPval takes its name from Gross Domestic Product: OpenAI started with GDP as its economic anchor and pulled tasks from the occupations that contribute most to it. The company described the effort as part of its mission "to ensure that artificial general intelligence benefits all of humanity."
The design marks a deliberate break from the benchmarks that have dominated AI evaluation. Academic tests like MMLU and competitive challenges sharpen model reasoning, OpenAI acknowledged, but they "often fall short of the kind of tasks that many people handle in their everyday work."
GDPval sits at the end of a progression OpenAI has been building for years: from MMLU's exam-style questions, to applied evaluations like SWE-Bench Verified (software bug fixes), MLE-Bench (machine learning engineering), and Paper-Bench (scientific critique of research papers), to market-priced evaluations like SWE-Lancer, which pays out based on real freelance software contracts.
What makes GDPval different is scope and realism. Unlike SWE-Lancer, which concentrates on software engineering, GDPval covers many occupations at once. Unlike MMLU or Humanity's Last Exam, which synthesize tasks in the style of academic tests, GDPval tasks are built from real work products: a legal brief, an engineering blueprint, a customer support conversation, a nursing care plan.
The tasks are not simple text prompts either. Each comes with reference files and context, and expected deliverables include documents, slides, diagrams, spreadsheets, and multimedia. That format, OpenAI argues, makes GDPval "a more realistic test of how models might support professionals."
How the occupations were chosen
The selection process was data-driven. OpenAI started with the 9 industries contributing over 5% of U.S. GDP per St. Louis Fed data. Within each industry, it picked the 5 occupations contributing most to total wages and compensation, using the May 2024 U.S. Bureau of Labor Statistics occupational employment report.
To filter for knowledge work specifically, OpenAI classified every task in O*NET — the U.S. Department of Labor's occupational database — as either knowledge work or physical labor. An occupation qualified as "predominantly knowledge work" if at least 60% of its component tasks involved no physical work. The company chose that threshold for the first version, targeting occupations "where AI could have the highest impact on real-world productivity."
The resulting 44 occupations run from software developers and lawyers to registered nurses and mechanical engineers.
Each task went through roughly 5 rounds of expert review, including checks from other task writers, additional occupational reviewers, and model-based validation. OpenAI says it deliberately recruited a breadth of experts — lawyers from different practice areas and firms of different sizes — to maximize representativeness. The dataset holds 30 fully reviewed tasks per occupation in the full set, and 5 per occupation in the open-source gold set.
Grading by blind comparison
Scoring relies on expert graders from the same occupations represented in the dataset. The graders blindly compare model-generated deliverables against those produced by the task writers, without knowing which is which, then rank them and classify each AI output as "better," "as good as," or "worse than" the human work. Task writers also built detailed scoring rubrics per occupation for consistency.
OpenAI additionally trained an "automated grader" — an AI system that predicts how human experts would judge a deliverable — and released it at evals.openai.com as an experimental research service. The company was blunt about its limits: it "isn't yet as reliable as expert graders," so OpenAI does not use it as a replacement.
The results: Claude Opus 4.1 leads, GPT-5 close behind
OpenAI ran blind evaluations across the 220 tasks in the gold set, comparing deliverables from GPT-4o, o4-mini, OpenAI o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4 against human-produced work.
Claude Opus 4.1 — Anthropic's flagship, not OpenAI's own — was the best performing model in the set. It produced outputs rated as good as or better than humans in just under half the tasks, and OpenAI credited it with excelling on aesthetics, such as document formatting and slide layout. GPT-5, in turn, stood out on accuracy, including finding domain-specific knowledge.
The trajectory matters as much as the leaderboard. Performance more than tripled from GPT-4o (spring 2024) to GPT-5 (summer 2025), a span OpenAI measures at just over a year, following what the company calls a clear linear trend.
The cost and speed figures are the headline economics. Frontier models completed GDPval tasks roughly 100x faster and 100x cheaper than industry experts. OpenAI attached a significant caveat: those numbers reflect pure model inference time and API billing rates, and "do not capture the human oversight, iteration, and integration steps required in real workplace settings." Still, the company concluded that on the subset of tasks where models are strongest, "giving a task to a model before trying it with a human would save time and money."
OpenAI also ran a training experiment. It incrementally trained an internal, experimental version of GPT-5 on GDPval-related data and found the process improved performance, which the company described as "a pathway for further potential improvement." Three other controlled experiments backed up the finding that capability on these tasks is tractable: increasing model size, encouraging more reasoning steps, and giving richer task context each produced measurable gains.
The full results appear in an accompanying paper, and OpenAI is releasing the gold subset of tasks and the public grading service so other researchers can build on the work.
What it means for work
OpenAI's own reading of the early data is measured. Models can already take on "some repetitive, well-specified tasks faster and at lower cost than experts," the company wrote. But it added that "most jobs are more than just a collection of tasks that can be written down," and framed GDPval as a map of where AI can absorb routine work "so people can spend more time on the creative, judgment-heavy parts of work." When AI complements workers this way, OpenAI argued, "it can translate into significant economic growth."
The benchmark's limitations are explicit. The current version is one-shot only, so it misses the reality that professionals iterate — revising a legal brief after client feedback, or re-running an analysis after spotting an anomaly. Real tasks also arrive without clean prompts and reference files; a lawyer may need to talk through ambiguity with a client before a brief is even the right deliverable. OpenAI says future versions will add interactivity, ambiguity navigation, and broader coverage of occupations and industries, with the long-term goal of better measuring progress on diverse knowledge work.
For now, GDPval offers something the AI field has lacked: a repeatable, expert-graded yardstick for whether models can do the work people are actually paid to do. OpenAI is inviting industry experts and enterprise customers to contribute to future rounds — and with performance tripling in roughly a year on this measure, the next iteration of the leaderboard will be watched closely by anyone tracking when AI assistance turns into AI substitution.
Original: bls.gov
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
119 articles
Related articles
- OpenAI Tells Business Leaders: Write Evals, Not Wish Lists
- OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
- OpenAI Maps 148 Million US Jobs Into Four AI Transition Paths
- OpenAI Lays Out Its Roadmap for the Next Phase of Enterprise AI
- OpenAI Says Agents Have Replaced Chatbots as Its Default Work Tool