
The classroom note
At Edo Boys High School in Benin City, Nigeria, first-year senior secondary students spent six weeks in an after-school programme (June–July 2024) built around Microsoft Copilot, powered by GPT-4, used as a virtual English tutor alongside their teachers. A World Bank blog post described the programme and previewed results the authors said were “soon to be published.” The formal Policy Research Working Paper, WPS11125, followed in May 2025, describing a randomised controlled trial run under the World Bank's Edo Basic Education Sector and Skills Transformation Operation. Pen-and-paper assessments covering English, AI knowledge and digital skills, plus end-of-year curricular exams and attendance monitoring, were the outcome measures.
What the evidence says
The working paper's own abstract reports a 0.31 standard deviation gain on the combined assessment and a 0.23 standard deviation gain on English specifically, the trial's primary outcome. The authors frame that as equivalent to roughly 1.5 to 2 years of “business-as-usual” schooling and, per the earlier blog post, among the top tier of education interventions in a comparative international database. A subgroup analysis found the largest effects among female students and those with stronger baseline performance, meaning the average result was not evenly shared across the sample. The blog post, written before the underlying paper was public, is promotional in tone; the working paper is the primary record and the more careful account of design and sample.
The implementation question
Six weeks, one school, one subject, one licensed product, delivered after school hours alongside participating teachers: that is a small and resource-intensive configuration to scale into an ordinary school day across a state system. The programme required device access, connectivity, a licensed AI product, and teacher time to run the sessions, none of which the published material prices out. A district reading this pilot needs to ask what it would cost to replicate the staffing and technology, not only whether the effect size looks large.
What holds and what fails
The result holds as an unusually strong six-week outcome from a randomised design at one school, worth taking seriously rather than dismissing. It does not yet establish that the gain persists after the programme ends, that it would hold in a different subject or school system, or that a shorter, cheaper, or less closely supported version would produce anything similar. The World Bank's own account calls this the first study of its kind in a developing-country context, a reason for caution about generalising, not for treating it as settled. This is an editorial point: a single early pilot, however striking, is not evidence of what a national rollout would achieve.
- Is this result from a randomised pilot, and over what period was it measured?
- What did the programme cost per student, including devices, connectivity, licensing and staff time?
- Did every subgroup benefit equally, or did the average conceal who gained least?
A six-week gain is real data. It becomes a policy the moment someone treats it as a promise about next year.
Sources & reading trail
Describes the after-school pilot, its six-week duration, outcome measures, and previews the reported effect size.
Source published: 9 January 2025 · Retrieved: 16 September 2026
Confirms the RCT design, the Microsoft Copilot/GPT-4 tool, and gives the exact standard-deviation results and subgroup findings.
Source published: 19 May 2025 · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.