ChatGPT and Critical-Thinking Training Boost Students in Different Ways
A randomized Bocconi experiment with 1,000+ students found ChatGPT access lifted rubric scores nearly a full point, while causal-reasoning training made ideas more original — and rubrics missed it.
Updated
Why it matters
- Over 1,000 first-year Bocconi University undergraduates participated in the randomized experiment, run with OpenAI Economic Research.
- Students with ChatGPT (GPT-4o) access scored almost a full point higher on a five-point grading rubric.
- Causal-reasoning training did not raise rubric scores but increased the variety and uniqueness of student ideas.
- Students who received both ChatGPT and the training showed gains across the widest range of measures.
- The rubric measured only two marketing goals — awareness and use of the university store — and missed the originality gains.
More than 1,000 first-year undergraduates at Bocconi University scored almost a full point higher on a five-point grading scale simply by having access to ChatGPT — while a separate critical-thinking exercise made their ideas measurably more original, not more polished.
That is the core finding of a randomized experiment run by researchers at Bocconi University in collaboration with OpenAI Economic Research. It is one of the cleaner attempts yet to answer two questions now facing every school and university: what does AI access actually do to student work, and does learning to think critically still matter when a chatbot can produce professional-sounding answers?
The study's answer is that both matter, but for different things — and that traditional grading may be blind to half the picture.
How the experiment worked
Students worked on a real-world business case: developing marketing recommendations for Bocconi's own merchandise store. The researchers randomly assigned entire class periods to one of four groups:
- Access to ChatGPT (GPT-4o)
- Training in causal reasoning, a specific form of critical thinking
- Both ChatGPT access and the training
- Neither
The causal-reasoning training had nothing to do with AI. It taught students to link cause and effect — to explain why a given solution might or might not work — through a game, examples, questions, and feedback.
Grading was two-track. Trained human graders scored submissions on a five-point rubric. Separately, automated text analysis measured each submission's number and variety of ideas, signs of causal reasoning, and similarity to submissions written by three experts.
The randomized design is what gives the study its weight. By separating the effects of ChatGPT access, the training, and their combination, the researchers could attribute each measured change to a specific intervention. That makes the experiment, as the researchers frame it, a useful contribution to a rapidly growing body of research on AI's impact on students and how to structure and support their learning.
What did ChatGPT access change?
Students with ChatGPT access produced better work by every conventional measure. Their answers included more ideas, followed clearer logic, and were more similar to recommendations written by experts. In short: AI helped novices produce work that looked more professional.
One detail matters for the debate over whether students simply outsource assignments to a chatbot. They did not hand the task over wholesale. As the researchers note, students still had to decide what to ask, evaluate the responses, and choose what went into their final submission. The polish came from AI; the selection and judgment remained with the student.
What did the critical-thinking training change?
Here the study produced its more unexpected result.
Students who completed the causal-reasoning exercise explained more clearly why their ideas might work and when they might fail. But they did not score higher on the grading rubric — which only measured how well recommendations addressed two standard marketing goals: increasing awareness of the university store and increasing its use.
The rubric missed something the text analysis caught. Across the group, students who completed the exercise produced a wider range of ideas, and those ideas were more distinct from what their peers produced. That gap is the study's sharpest finding about assessment itself: a traditional rubric can reward a clear, well-structured answer while overlooking whether a student came up with an idea no one else did.
What happened when students got both?
The combined group showed the benefits of each intervention, and the broadest gains overall.
- Their idea variety matched that of students who completed only the exercise.
- Their rubric scores and number of ideas matched those of students with ChatGPT access alone.
- Their work showed stronger logical coherence and more evidence of searching for explanations and questioning assumptions.
The researchers are explicit about the framing this supports. It can be tempting to cast the education debate as a choice between learning to think for yourself and learning to use AI. This experiment says both are valuable in different ways, and together they are complementary.
As the study puts it, the takeaway is: "AI helped students make their answers better. Critical-thinking training helped make their ideas broader."
Why the assessment gap matters
The mismatch between what the causal-reasoning exercise did for students and what the traditional rubric captured points to a broader challenge for schools.
If AI can help students produce polished, expert-like work, then the final answer alone tells educators less about what a student actually understands. The polished artifact is no longer reliable evidence of the underlying skill.
Many educators are already grappling with the implication that assignments and evaluations need to change. The researchers situate this in a multi-decade pattern of technology and education evolving together in service of the skills students need in modern society.
The concrete recommendation that follows from the data: reward students for producing work that reflects originality, reasoning, and consideration of multiple approaches — not just the most conventional or polished answers.
The stakes beyond one classroom
The study lands amid an international argument over AI in education, with schools split between restriction and integration. Its contribution is to replace a binary question with two measurable, separable effects. AI access raised the quality and coherence of student work on a real assignment. Critical-thinking training — delivered without any AI component — increased the uniqueness of student ideas.
Neither crowded out the other. Students who received both showed gains across the widest range of measures in the experiment, which suggests the practical path for institutions is not choosing between AI literacy and thinking skills, but sequencing and combining them — while redesigning the rubrics used to judge the result.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
200 articles