RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Evidence

Evidence / Reference note · Reference note · prepared 16 September 2026

A rigorous trial and a school pilot measure different things

A funder's evidence tiers, a Harvard RCT and an EEF school trial show what a genuine trial needs that a one-term pilot usually cannot supply.

Visual for this record: A rigorous trial and a school pilot measure different things
Visual published by ies.ed.gov, shown for identification of the record. Credit: ies.ed.gov · source page ↗ Rights: owner-review-pending.

The classroom note

A department running a one-term pilot wants to know, at the end, whether it "worked." What that question can answer depends on decisions made before the pilot starts: what is measured, against what comparison, judged by whom. Reading how a formal evidence standard, a university study and a school trial answered those questions, as retrieved on 16 September 2026, shows how far a single-term pilot sits from any of them.

What the evidence says

The What Works Clearinghouse's ESSA page sets a formal floor for "strong" or "moderate" evidence: a statistically significant positive effect, "at least 350 students," and "at least two educational sites." A randomised controlled study at Harvard cleared a version of that bar for one outcome: 194 students in the university's largest introductory physics course were split into comparable groups, each experiencing both a custom AI tutor and active-learning teaching across two weeks, with learning gains measured on a common instrument. Even so, its authors caution they "do not presume" the AI tutor will outperform active learning "in all contexts," naming complex synthesis and critical thinking as likely exceptions. The EEF's "Teacher Choices" trial on ChatGPT-assisted lesson planning took a narrower, school-based approach: 68 schools, an independent evaluator (NFER), and one pre-specified outcome, finding teachers using ChatGPT spent 56.2 minutes a week on lesson preparation against 81.5 minutes in a comparison group, with quality checked by an expert panel reviewing work blind to which method produced it. EEF's own rating system records no formal evidence-strength score for this trial type.

The implementation question

What separates all three from an ordinary pilot is structure, not effort: a comparison group experiencing something other than the new tool, one outcome decided in advance rather than assembled afterward from favourable anecdotes, and someone judging results without knowing which group is which. A school piloting a tool for a term can borrow that structure at a smaller scale - track one measurable thing against a comparable class or previous term, rather than gathering general impressions once the pilot ends.

What holds and what fails

What holds: a well-designed pilot, even a small one, can tell a school whether a tool merits further resources. What fails: treating a single term in one school as evidence of a general effect, when the cited standard requires hundreds of students across multiple sites and the university study's own authors decline to generalise beyond their subject and format. This is an editorial framework built from these three documents, not a fixed procedure; any pilot design should still fit a school's own capacity and ethics requirements.

  • What single outcome are we measuring, decided before the pilot starts rather than chosen from whatever looks favourable afterward?
  • What is this being compared against - a similar class, a previous term, nothing at all?
  • Who is judging the result, and could they do so without knowing which group used the tool?

A pilot answers a local question about one school's own conditions. Only a larger, replicated trial can answer the general one - and even then, as the Harvard study shows, not for every context.

Sources & reading trail

WWC | ESSA Tiers Of Evidence ↗

Sets the formal sample (350 students) and site (two sites) floor and significance requirement behind strong/moderate ESSA evidence ratings.

Source published: Not established · Retrieved: 16 September 2026

AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting ↗

Randomised controlled trial (N=194, Harvard, Fall 2023) comparing a custom AI tutor with active-learning teaching; authors caution against generalising beyond their subject and format.

Source published: 3 June 2025 · Retrieved: 16 September 2026

ChatGPT in lesson preparation - Teacher Choices trial ↗

68-school EEF Teacher Choices trial with an independent evaluator (NFER) found ChatGPT-assisted lesson planning saved teachers 25.3 minutes a week versus a comparison group, with quality checked by blind expert review; completed September 2026, no formal evidence-strength rating assigned.

Source published: Not established · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.