
The classroom note
At a large high school in Turkey, maths classes across roughly 50 classrooms were split for a term into three arms: students with no AI access, students given a plain ChatGPT-style interface built on GPT-4 (“GPT Base”), and students given the same model wrapped in prompts supplying teacher-designed hints instead of answers (“GPT Tutor”). The randomised field experiment, published in the Proceedings of the National Academy of Sciences, involved nearly 1,000 students in grades 9 to 11 during the 2023-2024 school year and was funded by the Wharton AI and Analytics Initiative, the Fishman-Davidson Center and Wharton Global Initiatives. During assisted practice, GPT Base raised scores by 48% and GPT Tutor by 127% against the no-access group. On a later exam sat without AI access, the picture reversed for one group.
What the evidence says
Students who had practised with GPT Base scored 17% lower on the unassisted exam than students with no AI access at all; GPT Tutor students showed no significant difference from control. The researchers, whose faculty page lists the same PNAS paper alongside a link to their own PDF, attribute the gap to how students used the tool: GPT Base's own answers were wrong in roughly half of responses, mostly through logical rather than arithmetic mistakes, and message logs showed students leaning on the model as what the authors call a “crutch” during practice rather than checking its output. Because the guardrail version withheld direct answers, it left students doing more of the reasoning themselves, which the authors argue is why its practice gains did not evaporate once the tool was removed.
The implementation question
The mechanism is a design choice inside the software, not a property of AI in general: does the interface hand over a finished answer, or does it structure a hint and require the student to complete the step. That single toggle, in this trial, decided whether access helped retention or quietly borrowed against it. A school evaluating any AI maths tool is really asking a vendor to show which of these two behaviours its product defaults to, and whether a teacher or student can change it.
What holds and what fails
The result holds as evidence that unrestricted access to a capable model can trade short-term practice performance for weaker independent recall, in this subject, this age group, this one school, over one term with an early-2024 version of GPT-4. It does not establish that every generative AI tool harms learning, or that guardrails always neutralise the effect; the authors themselves caution that findings may not generalise to subjects without objective evaluation criteria, and that a newer model might behave differently.
- Does this tool give answers, or does it require the student to produce the next step themselves?
- Is the reported gain measured with AI assistance present, without it, or both?
- What happens to performance when the tool is taken away, and has anyone checked?
A percentage gain recorded mid-practice and a percentage change recorded afterwards are two different facts, and this trial is useful mainly for showing how far apart they can be.
Sources & reading trail
Gives the RCT design, three treatment arms, sample of nearly 1,000 students, funder, practice/exam results and stated limitations.
Source published: 25 June 2025 · Retrieved: 16 September 2026
Confirms authorship and venue of the study and links the author's own copy of the paper.
Source published: Not established · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.