Google's Gemini Tutor Boosted Sierra Leone Math Scores 0.258 SD
Google's Guided Learning in Gemini produced a +0.258 standard deviation math gain for 1,763 Sierra Leone students over 8 weeks; heavy-use classrooms gained the equivalent of 1.8 to 2.5 years of typical learning progress.
Updated
Why it matters
- +0.258 standard deviation math gain for 1,763 students across 12 schools over 8 weeks in Sierra Leone's Port Loko District
- Heavy-use classrooms (12 hours of Gemini use / ~half of math lessons) saw 1.8 to 2.5 years of typical learning progress
- 69% of students met or exceeded usage targets, versus ~5% typical for voluntary ed-tech (the 'Five Percent Problem')
- Analysis of 113,000+ interactions found 91.4% focused on conceptual understanding, with Gemini giving direct solutions in only 2% of cases
- Skill-building queries rose from 68% in week 1 to 90% by week 8, while solution-seeking fell from 25% to 10%
Google's Guided Learning mode in Gemini produced a +0.258 standard deviation gain in math scores for 1,763 junior secondary students across 12 schools in Sierra Leone's Port Loko District, according to a randomized controlled trial published this week. The effect translates to 1.2 to 1.7 years of typical learning progress compressed into an eight-week window — and up to 2.5 years for classrooms that used the tool most heavily.
The study, run in partnership with Fab AI and the Sierra Leone Ministry of Education with funding from Google.org and the Gates Foundation, is among the largest field trials of a commercial generative-AI tutor in a low-income classroom setting. Google's report frames the result as evidence that "AI can be a powerful pedagogical partner — not by replacing teachers, but by augmenting their reach."
What did the trial actually measure?
Researchers randomly assigned classrooms to either use Gemini's Guided Learning mode alongside standard instruction or follow the existing curriculum. The trial ran for eight weeks during the school term. Researchers pre-registered the study, filing its methods and outcome measures publicly before collecting results.
The team captured more than 113,000 student-AI interactions during the trial and coded each one to determine whether students used the tool to build understanding or to extract direct answers. 91.4% of conversations focused on conceptual understanding. Gemini responded with scaffolding questions in 76% of its messages and produced direct solutions in only 2% of cases.
That posture is built into the product. Guided Learning emerged from Google's LearnLM research effort and is, per the report, "pedagogically-grounded and specifically tuned to prioritize building understanding over providing direct answers."
How large was the learning gain?
The headline number — +0.258 standard deviations — sits at the upper end of effect sizes commonly seen in education research, where most interventions register closer to +0.10 SD. In practical terms, the report estimates the gain as 1.2 to 1.7 years of typical math progress in eight weeks.
The effect grew with exposure. Classrooms whose teachers hit roughly 12 hours of Gemini use — about half of math lessons during the trial — saw students post gains equivalent to 1.8 to 2.5 years of typical learning progress. The engagement figure is one of the more striking results: 69% of students met or exceeded usage targets, a number Google contrasts with the roughly 5% participation rate considered normal for voluntary educational technology — what the company calls "The Five Percent Problem."
"That means students were not only engaged but they enjoyed coming to class more," the report states.
What did teachers and students do differently?
Teachers remained the architects of the classroom. They set lesson objectives, designed activities, and used Gemini primarily as a discussion partner and lesson-planning aid. In focus groups, "many described a shift from 'lecturers' to 'facilitators,' moving through the classroom to support pairs of students as they navigated their own learning journeys," the report says. Several teachers said the tool helped them find new ways to explain familiar topics like fractions.
Students also changed how they used the AI. Skill-building queries rose from 68% of conversations in week one to 90% by week eight. Solution-seeking dropped from 25% to 10% over the same period — what the report describes as proof that "students didn't just want answers, they wanted to understand how they got there."
That behavioral shift cuts against a common critique of generative AI in schools: that chatbot tutors erode critical thinking by handing out answers. The Sierra Leone data, at least in this setting, points the other way — students trained the model, and themselves, toward more rigorous use.
Where did the model fall short?
The study surfaced an achievement gap that the report flags as the central design challenge ahead. While most students benefited, those who entered the trial with stronger math skills captured the largest gains. The pattern mirrors a recurring concern in ed-tech: tools that personalize well often personalize best for students who already have momentum.
Google says it plans to address this directly in subsequent trials by studying "metacognition and relational intelligence" — the cognitive habits and social dynamics that drive independent learning — to "capture a more holistic view that explores the nuanced complexity of learning."
Who ran the study, and what is being released?
Google led the trial in partnership with Fab AI and the Sierra Leone Ministry of Education. Google.org and the Gates Foundation funded the work; EducAid, Laterite, and Oxford MeasurEd contributed research support.
Google released two open documents alongside the results. A teacher training guide includes the specific protocols used in the trial, and an RCT playbook documents the team's approach to running pre-registered studies in low-resource settings — aimed at helping other researchers "run faster, scalable studies aligned to their needs and contexts."
The work also feeds into the Global AI for Learning Alliance (GAILA), a multi-stakeholder effort to coordinate evidence on AI in education across countries.
Why does this trial matter outside Sierra Leone?
Generative AI tutors have, until now, been evaluated mostly in short demonstrations or in higher-income school systems. The Sierra Leone study offers a rare counterpoint: a pre-registered, classroom-level RCT in a low-bandwidth environment, run at meaningful scale, with behavior-coded interaction logs.
The next round of trials, Google says, will probe metacognition and relational intelligence and run in additional countries to build "a more comprehensive, cross-country evidence base, which we hope will inform responsible development of AI across the learning ecosystem." Whether the Sierra Leone effect holds across languages, curricula, and longer time horizons will determine whether the result reads as a proof point for AI tutors — or as a reminder that eight weeks is not a school year.
Original: storage.googleapis.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
237 articles
Related articles
- OpenAI's Study Mode Boosted Exam Scores Up to 15%
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- Google DeepMind Launches ATL Saathi, a Gemini Assistant for Indian Teachers
- OpenAI Brings AI Skills Jam to 1,600 K–12 Educators Across Eight US Cities
- Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel