
The classroom note
In Harvard's introductory Physical Sciences 2 course, students spent alternating weeks either working through material with a purpose-built AI tutor or taking part in the course's normal active-learning sessions, the format most physics-education research has favoured for a generation. A randomised crossover trial published in Scientific Reports reports that 194 students met the study's inclusion criteria out of 233 originally enrolled, and each experienced both conditions in turn. On a post-test given after each unit, the AI-tutored group's median score of 4.5 significantly exceeded the active-learning group's median of 3.5, against a shared pre-unit baseline of 2.75. The paper's five authors, led by Greg Kestin, report effect sizes of 0.73 to 1.3 standard deviations, describing the gap as more than double the active-learning group's gain over baseline.
What the evidence says
A crossover design lets each student serve as their own control, strengthening the comparison inside this single course. The open-access copy hosted on PubMed Central adds detail the journal summary omits: the tutor ran on GPT-4 behind expert-crafted, question-specific prompts instructors wrote in advance, and the paper carries no external funding statement and no competing interests declared. Students using the tutor also spent less time on task, a median 49 minutes against roughly 60, and reported higher engagement (4.1 versus 3.6) and motivation (3.4 versus 3.1) on five-point scales. The authors are explicit that their material targeted understanding, application and early analysis, not the complex synthesis later coursework demands, and that the result describes a pedagogically engineered tutor, not an unguided chatbot a student might open alone.
The implementation question
The mechanism worth naming is the engineering, not the model. Building this tutor required instructors to write step-by-step solutions and question-specific prompts before a student used it, then maintain that scaffolding across a term. That is a cost in staff time and subject expertise, concentrated at one well-resourced university department, before any classroom benefit appears. A school considering something similar is really asking whether it can fund that authoring work, not whether a general-purpose chatbot will reproduce the result.
What holds and what fails
The finding holds as evidence that a carefully authored AI tutor can match or exceed one specific active-learning format on one kind of learning outcome, in one course, at one institution. It fails as evidence for AI tutoring in general, for other subjects, for younger pupils, or for open-ended AI use without the instructor-written scaffolding the authors built. Treating a single-course crossover trial as proof that AI tutoring works is, editorially, exactly the overreach the authors themselves warn against.
- Does the tool in front of us include expert-authored, subject-specific prompting, or is it a general chatbot left to improvise?
- Is the claimed gain measured on the same kind of task the trial tested, or is it being stretched to cover judgement and synthesis?
- Who will write and maintain the subject-specific scaffolding here, and is that cost accounted for?
A well-run trial in one lecture hall is real evidence, not a verdict on every classroom that might buy a similarly branded product.
Sources & reading trail
Gives the randomised crossover design, sample of 194 students, and the headline post-test and effect-size results.
Source published: 3 June 2025 · Retrieved: 16 September 2026
Adds the funding/competing-interests statements, the GPT-4 prompt-engineering description, and the authors' stated scope limits.
Source published: 3 June 2025 · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.