RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Classroom practice

Classroom practice / Reference note · Reference note · prepared 16 September 2026

A guardrail, not the chatbot, made a maths AI tutor safe

A Turkish RCT found an unguarded chatbot cut later exam scores while a hint-only version did not, unlike a vendor's design claims.

pnas.orgprimary record

Generative AI without guardrails can harm learning: Evidence from high school mathematics

Document
25 June 2025
Event
no single event
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The classroom note

Roughly fifty classes of ninth- to eleventh-graders in a Turkish high school worked through the same maths practice questions, some with a laptop and a chat window, some without. That is the setting for a randomised trial that gives this argument something rare: a measured comparison, not a marketing claim, of what changes when a student practises with a generative AI tutor rather than a textbook. It matters because products a family is more likely to meet, Khan Academy's Khanmigo chief among them, describe their own design in language close to the one the trial found protective, without publishing an equivalent independent test.

What the evidence says

The preregistered randomised controlled trial, run with nearly 1,000 students in Fall 2023, assigned each class to a no-AI control, a plain ChatGPT-style tool (‘GPT Base’), or one prompted to give hints rather than answers (‘GPT Tutor’). On assisted practice, both AI arms scored higher than the control, 48% and 127%. On a later closed-book exam with no AI access, GPT Base students scored 17% worse than the control, while GPT Tutor was statistically indistinguishable from it. The authors traced the harm to GPT Base's frequent wrong answers, correct only 51% of the time, and to students using it as a ‘crutch’ to copy solutions. They are explicit about scope: one school, one subject with gradable answers, one semester.

The implementation question

Khanmigo's page describes close to the same guardrail the trial found protective: it ‘doesn't just give answers’ but ‘guides learners to find the answer themselves,’ citing a four-star Common Sense Media rating. Neither is a measurement of unassisted performance after use; the rating assesses design and safety, not outcomes, and a teacher's quoted account of a rubric task shrinking from an hour to fifteen minutes is one anecdote about the teacher's time, not the pupil's exam score. Access is gated: families pay for the learner product, and classroom access runs only through a paid school or district contract, a cost worth separating from any evidence of what the minutes saved bought.

What holds and what fails

What holds, on the trial's evidence, is the mechanism: a tutor built to withhold the answer and hint instead preserves practice gains without the exam-score cost a plain answer-giving chatbot produced. What fails is assuming any product using guardrail language has been tested the way this trial tested a purpose-built tool; Khanmigo's description matches the protective arm, but the cited study evaluated a separate, academically built system, not Khanmigo itself.

  • Does the product report unassisted performance after use, or only performance while the tool is available?
  • Is ‘guides rather than answers’ a tested behaviour or a design description?
  • What does a family or school pay for access, and does that cost track any independent measurement of learning?

A tutor that hints and a tutor that answers can look identical in a demonstration and produce opposite results in an unassisted exam; the trial's contribution is showing that the difference is measurable, not just describable.

Sources & reading trail

Generative AI without guardrails can harm learning: Evidence from high school mathematics ↗

Preregistered RCT design, sample and the practice-score gain versus unassisted-exam-score loss for an unguarded chatbot versus a hint-only tutor.

Source published: 25 June 2025 · Retrieved: 16 September 2026

Khanmigo: Khan Academy's AI-powered teaching assistant & tutor ↗

Vendor's own description of Khanmigo's answer-withholding design, a cited third-party rating, a teacher testimonial, and access/pricing terms.

Source published: Not established · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.