
The classroom note
Khan Academy launched Khanmigo, its generative AI tutor and teaching assistant, in 2023, and by 2026 the organisation has begun publishing figures on how the tool performs. A blog post dated 6 May 2026 reports internal testing across roughly 15 million tutoring threads over six months, with experiments drawing on samples of 352,000 to 1.36 million threads. The organisation reports a 6.1% improvement in “next-item correctness,” meaning whether a student got the next practice question right, a 5.09% increase in “cognitive engagement quality,” and a 50% reduction in instances where the tutor gave away an answer rather than guiding the student to it.
What the evidence says
These figures come from Khan Academy's own internal A/B testing, using its own metrics, not an independent researcher or a randomised trial with an external comparison group of the kind found in peer-reviewed studies of other AI tutors. The post frames the work as measuring “independent learning transfer” rather than performance with the tutor's help present, a meaningful choice, but reported by the vendor rather than verified by an outside evaluator. A second post, published 15 June 2026, names three pilot districts, Hanover Community School Corporation in Indiana, Taft Independent School District in rural Texas, and Christopher Columbus High School, as “co-builders” who shaped product design rather than subjects of formal evaluation, and reports district students were six times more likely to reach recommended practice levels than independent learners, a usage measure rather than a learning-outcome measure.
The implementation question
A district asking does Khanmigo work is really asking two separable questions Khan Academy's material does not fully answer: does the tutor improve engagement and platform usage, which the company's data speaks to directly, and does it improve learning outcomes independent of practice volume, which requires a study design the company does not claim to have run. The company's own text is candid here, stating “practice is not the whole story of learning” and that teachers, curriculum and school context all matter alongside any tool.
What holds and what fails
The reported figures hold as evidence of what Khan Academy's telemetry shows about engagement and short-term practice behaviour at large scale, which is not nothing. They do not hold as independent, peer-reviewed evidence of learning gains, since no third-party research partner, control group, or published methodology comparable to the randomised trials run on other AI tutors appears in this material as retrieved. Vendor-reported figures are not disqualified by being vendor-reported, but they carry a different evidentiary weight than the Harvard and Nigeria trials this site has covered, and should be labelled as such.
- Is this figure from the vendor's own telemetry, or an independent researcher with no stake in the product?
- Does the measure describe engagement and practice volume, or a learning outcome checked against a comparison group?
- What would an independent, peer-reviewed evaluation of this tool need to look like before we treat it as settled?
A company that publishes its own numbers is doing more than one that publishes none, but a school still has to ask who checked the number before repeating it to a parent.
Sources & reading trail
Gives Khan Academy's self-reported internal A/B testing scale and metrics for Khanmigo's tutoring behaviour.
Source published: 6 May 2026 · Retrieved: 16 September 2026
Names pilot districts and frames their involvement as product co-development rather than independent evaluation, with a usage statistic.
Source published: 15 June 2026 · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.