
The classroom note
A science teacher spending an evening on lesson preparation is exactly the workload the Education Endowment Foundation's ChatGPT in Lesson Preparation trial set out to measure. Sixty-eight state-funded secondary schools in England and 259 teachers of Year 7 and 8 science were randomly allocated, at school level, to a ChatGPT group or a group asked not to use any generative AI, with independent evaluator NFER running the trial and the EEF publishing its evaluation report on 12 December 2024. ChatGPT teachers had a five-week period to learn the tool, using a guide built by Bain & Company's Social Impact practice and co-funded by the Hg Foundation, before planning time in weeks six to ten was logged in an online diary.
What the evidence says
This was a genuine cluster-randomised controlled trial, not a comparison of volunteers who chose ChatGPT: an NFER statistician randomly allocated whole schools to each arm after baseline surveys closed on 19 March 2024. ChatGPT-group teachers recorded a median 56.2 minutes a week on preparation against 81.5 minutes in the comparison group, a saving of 25.3 minutes, or 31%, which the EEF rates a high-security finding. A panel blind to group allocation reviewed submitted lesson resources and found no discernible quality difference. The report also found the share of ChatGPT-group teachers who felt they spent “too much time” on preparation fell from 49% to 26% during the trial, with no comparable shift in the non-AI group.
The implementation question
The trial measured teacher time and teacher-reported workload; it did not measure whether pupils in either group learned more, less, or the same. A district reading “31% less planning time” as evidence that ChatGPT improves teaching is filling a gap the trial itself left open. The mechanism was narrow, too: teachers mostly used ChatGPT for one or two activities, commonly creating quizzes or generating activity ideas, not rewriting entire lessons.
What holds and what fails
The time saving holds as a well-powered, randomised, independently evaluated result for the population who volunteered: NFER notes these teachers were likely more positively disposed to ChatGPT than teachers generally, and cites Teacher Tapp data suggesting the group who have never tried GenAI is shrinking. The report also flags a limitation: teachers were not blind to their own group while self-reporting time, and the ChatGPT group was asked to submit only ChatGPT-assisted resources, which could have flattered the quality comparison. It fails as evidence about pupil outcomes, other subjects or key stages, or other AI tools, none of which this trial tested.
- Is a reported saving about teacher time, workload, or pupil learning, and which does our decision depend on?
- Were participants randomly assigned, or did they opt in to the group whose result we are being shown?
- What tool, subject and year group was actually tested, and does it match what we intend to buy?
Less time on a task is a real, measurable good for a stretched profession; it is a different claim from a better lesson, and this trial keeps the two apart even where its headline invites readers not to.
Sources & reading trail
Gives the trial's design, sample of 68 schools and 259 teachers, evaluator, funder, and headline time-saving and quality results.
Source published: 12 December 2024 · Retrieved: 16 September 2026
Confirms cluster-randomisation at school level and gives NFER's own stated limitations, including self-report bias and submission self-selection.
Source published: 12 December 2024 · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.