RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Evidence

Evidence / From the archive · 6 May 2026 event · prepared 16 September 2026

Khan Academy's own data on its AI tutor lacks outside evaluation

Khan Academy's published Khanmigo results are internal engagement metrics and pilot-district accounts, not independent, peer-reviewed evaluation.

Visual published with the cited source for this record: Khan Academy's own data on its AI tutor lacks outside evaluation
Visual published with the cited source, shown for identification of the record. Credit: blog.khanacademy.org · source page ↗ Rights: owner-review-pending.

The classroom note

Khan Academy launched Khanmigo, its generative AI tutor and teaching assistant, in 2023, and by 2026 the organisation has begun publishing figures on how the tool performs. A blog post dated 6 May 2026 reports internal testing across roughly 15 million tutoring threads over six months, with experiments drawing on samples of 352,000 to 1.36 million threads. The organisation reports a 6.1% improvement in “next-item correctness,” meaning whether a student got the next practice question right, a 5.09% increase in “cognitive engagement quality,” and a 50% reduction in instances where the tutor gave away an answer rather than guiding the student to it.

What the evidence says

These figures come from Khan Academy's own internal A/B testing, using its own metrics, not an independent researcher or a randomised trial with an external comparison group of the kind found in peer-reviewed studies of other AI tutors. The post frames the work as measuring “independent learning transfer” rather than performance with the tutor's help present, a meaningful choice, but reported by the vendor rather than verified by an outside evaluator. A second post, published 15 June 2026, names three pilot districts, Hanover Community School Corporation in Indiana, Taft Independent School District in rural Texas, and Christopher Columbus High School, as “co-builders” who shaped product design rather than subjects of formal evaluation, and reports district students were six times more likely to reach recommended practice levels than independent learners, a usage measure rather than a learning-outcome measure.

The implementation question

A district asking does Khanmigo work is really asking two separable questions Khan Academy's material does not fully answer: does the tutor improve engagement and platform usage, which the company's data speaks to directly, and does it improve learning outcomes independent of practice volume, which requires a study design the company does not claim to have run. The company's own text is candid here, stating “practice is not the whole story of learning” and that teachers, curriculum and school context all matter alongside any tool.

What holds and what fails

The reported figures hold as evidence of what Khan Academy's telemetry shows about engagement and short-term practice behaviour at large scale, which is not nothing. They do not hold as independent, peer-reviewed evidence of learning gains, since no third-party research partner, control group, or published methodology comparable to the randomised trials run on other AI tutors appears in this material as retrieved. Vendor-reported figures are not disqualified by being vendor-reported, but they carry a different evidentiary weight than the Harvard and Nigeria trials this site has covered, and should be labelled as such.

  • Is this figure from the vendor's own telemetry, or an independent researcher with no stake in the product?
  • Does the measure describe engagement and practice volume, or a learning outcome checked against a comparison group?
  • What would an independent, peer-reviewed evaluation of this tool need to look like before we treat it as settled?

A company that publishes its own numbers is doing more than one that publishes none, but a school still has to ask who checked the number before repeating it to a parent.

Sources & reading trail

How Khan Academy Is Building a Better AI Tutor: Our Most Recent Learnings ↗

Gives Khan Academy's self-reported internal A/B testing scale and metrics for Khanmigo's tutoring behaviour.

Source published: 6 May 2026 · Retrieved: 16 September 2026

Built in the Open: How Pilot Districts Shaped the Reimagined Khan Academy ↗

Names pilot districts and frames their involvement as product co-development rather than independent evaluation, with a usage statistic.

Source published: 15 June 2026 · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.