OpenAI's Study Mode Boosted Exam Scores Up to 15%
A randomized trial of 300+ students found ChatGPT study mode lifted microeconomics exam scores ~15%. Now OpenAI is testing gains across 20,000 Estonian students.

Updated
Why it matters
- OpenAI's randomized study of 300+ college students found study mode access produced roughly 15% higher microeconomics exam scores versus a no-AI control; neuroscience gains were not statistically distinguishable from control.
- OpenAI built the Learning Outcomes Measurement Suite with Estonia's University of Tartu and Stanford's SCALE Initiative to longitudinally measure AI's impact on learning.
- The suite is being validated in Estonia with nearly 20,000 students aged 16-18 over several months before OpenAI releases it as a public resource.
College students with access to ChatGPT's study mode scored roughly 15% higher on microeconomics exams than peers studying with traditional online resources, according to a new OpenAI randomized controlled trial of more than 300 students. The company is now building a longitudinal measurement framework to answer the harder question its own research raised: do those gains persist over time?
The study, which OpenAI ran with students preparing for neuroscience and microeconomics exams, assigned participants to one of three groups. A control group studied with Google Search and YouTube, with AI-generated overview features disabled. Two additional groups received access to one of two study mode variants designed to guide students through the learning process in slightly different ways. OpenAI collected baseline quizzes and onboarding surveys to adjust for differences in prior coursework, study habits, academic confidence, and familiarity with AI tools.
The results were not uniform. In neuroscience, OpenAI observed directionally positive differences for study mode relative to control, but the results were not distinguishable from students studying with traditional online resources; the company said onboarding and technical issues cut into time spent studying among study mode users. In microeconomics, the company reported meaningful gains — roughly 15% higher scores relative to the no-AI control group, a result that held when each study mode variant was compared separately with the control.
OpenAI framed the study as an intention-to-treat analysis: the impact of being offered access to the tool under realistic rollout conditions, not the impact on students who used it intensively. Participation was not tied to exam performance, and not all students used study mode to the same extent during the nominal 40-minute sessions. The design reflected real-world study conditions rather than a tightly controlled lab environment, with the two study mode variants counterbalanced across subjects.
The stakes are considerable. Education is one of AI's most promising commercial and social frontiers — OpenAI notes that tools like ChatGPT can make personalized learning support available to any student, anywhere, at any time. But the education sector remains early in understanding how AI actually affects learning outcomes, and schools and policymakers worldwide are making adoption decisions with thin evidence. OpenAI's own analysis is still underway, but the company says early results give it confidence that a pedagogically aligned AI interaction style, encouraged through features like study mode, can improve learning outcomes.
Study mode, which OpenAI introduced last year, runs on custom system instructions written in collaboration with teachers, scientists, and pedagogy experts. The instructions are designed to support what OpenAI calls true learning rather than just answers: scaffolding, checks for understanding, and guided practice. Students using AI to study can mean anything from seeking quick answers to working through problems step by step with tutor-like guidance, and study mode pushes ChatGPT toward the latter behavior.
The measurement gap
The trial surfaced what OpenAI calls a deeper limitation in how learning outcomes are typically measured. Most existing evaluation approaches rely on fixed interventions assessed over short time windows, using test scores or final essays as primary signals. Those methods are not designed to capture the mechanism through which AI affects learning in practice: ongoing, personalized interactions that evolve alongside a learner's strategies, preferences, and study habits. They also fail to surface whether improvements in one capability, such as short-term recall, come alongside trade-offs in others, such as persistence, autonomous motivation, or creative problem solving.
"Most research methods focus on narrow performance signals — such as test scores — and lack the ability to assess how students actually learn with AI in real-world settings, and how that use shapes outcomes over time," OpenAI wrote in its announcement.
Because learning environments differ widely across countries, curricula, and institutional goals, OpenAI argues, outcomes from one-off studies rarely generalize across systems. Measurement approaches must be flexible enough for different education systems to define success in their own context.
The Learning Outcomes Measurement Suite
To address the gap, OpenAI developed the Learning Outcomes Measurement Suite, a framework built with Estonia's University of Tartu and the SCALE Initiative at the Stanford Accelerator for Learning to support longitudinal measurement of learning outcomes across different educational contexts.
The suite rests on three signals: how the model behaves, how learners respond, and what measurable cognitive outcomes result over time. It has five components.
System instructions to refine model behavior use natural language to change the default behavior of the model to align it with specific pedagogical approaches.
Learning interaction classifiers automatically detect "learning moments" within real, de-identified learner-model interactions and label characteristics such as engagement and error correction.
Learning quality graders score each learning moment by whether the learner achieved their objective and how closely the interaction followed strong pedagogical principles, including identification of failure modes.
Longitudinal learning graders track changes in the same learner's interactions with the model over time — engagement, persistence, and metacognitive strategies — at the individual and cohort levels.
Standardized cognitive and metacognitive measures are validated third-party instruments delivered via ChatGPT before, during, and after access, establishing baselines and measuring changes in capabilities such as critical thinking, creativity, and memory.
Combined, the system produces structured views of learning moments, dashboards showing how outcomes shift over time across cohorts, indicators of model performance against teaching and tutoring rubrics, and outcome measures aligned to standardized assessments and short learner questionnaires. Where available, it can incorporate partner-provided ground truth such as exam scores, classroom observations, or attendance. OpenAI says all data is de-identified.
The suite also tracks what OpenAI describes as holistic capabilities that underpin learning rather than narrow test-score definitions. These include autonomous motivation — whether learners shape their own studies rather than being directed by the model; productive engagement, measured by the frequency, variety, and quality of pedagogical interactions; task persistence — whether a learner pushes through cognitive challenges; metacognition, or the frequency and quality of planning, reflection, and monitoring; and recall of content from previous interactions.
OpenAI is explicit that it does not expect a single optimization target. "There will be no silver bullet in terms of what to optimize for," the company wrote. "Systems and educators will need to be empowered to guide trade-offs in alignment with pedagogical best practice and approaches."
Validation at nation scale
OpenAI is validating the suite through a large-scale randomized controlled trial before making it broadly available. The work is underway with the University of Tartu and Stanford's SCALE Initiative across nation-scale partners, including Estonia, where the measurement suite is being studied with nearly 20,000 students aged 16-18 over several months. Student use will happen in close collaboration with local leaders, OpenAI says, to ensure safety and alignment with local curricula.
"Estonia has always approached education not as static but as a system we continuously improve. With AI becoming part of that picture, the big question is how we measure AI's long-term impact on learning. That's what we're figuring out in collaboration with OpenAI," said Jaan Aru, Associate Professor at the Institute of Computer Science, University of Tartu. "Students are keen to be involved in the development process, and many want to learn how to support learning with AI. It feels like a real turning point, and we're excited to contribute methods that other education systems can reuse and build on."
Further research is planned through the founding organizations of OpenAI's Learning Lab, the company's new learning research ecosystem, which includes researchers from Arizona State University, UCL Knowledge Lab, and MIT Media Lab, building on prior collaborative studies. OpenAI is also supporting studies at the intersection of learning and labor — examining how AI shapes students' academic pathways, career decisions, and how institutions can support responsible adoption — at Bocconi University, Innova Schools, the Tuck School of Business at Dartmouth, San Diego State University, Stony Brook University, and others.
The framework has academic backing from Stanford. "This research allows us to learn quickly while also laying the groundwork for a deeper understanding of how AI can be thoughtfully integrated into schools in ways that truly matter," said Susanna Loeb, Professor of Education and Faculty Director of the SCALE Initiative at Stanford University. "We want to understand how these tools can support rigorous academic learning while also cultivating higher-order thinking, creativity, curiosity, and students' confidence in themselves as learners."
OpenAI says it intends to publish more research over time and release the measurement suite as a public resource for schools, universities, and education systems worldwide. For a company racing to embed ChatGPT into classrooms globally, the credibility of that evidence base — and whether the microeconomics gains hold up across months of use rather than a single exam cycle — will largely determine how education systems respond.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles
Related articles
- OpenAI Launches Learning Accelerator, a $500,000-Backed India-First Education Push
- OpenAI Launches First Certification Courses Inside ChatGPT
- OpenAI Says a Quarter of U.S. Workers Now Use ChatGPT on the Job
- OpenAI launches ChatGPT agent, folding Operator into ChatGPT
- OpenAI Signals Data Shows ChatGPT Use Deepening Worldwide