RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Evidence

Evidence / From the archive · November 2014 event · prepared 16 September 2026

Two meta-analyses put tutoring software near human tutors

Pooled effect sizes from 2014 and 2016 show intelligent tutors matching tutors on narrow tasks, not generative ones.

api.openalex.orgprimary record

Intelligent tutoring systems and learning outcomes: A meta-analysis

Document
1 November 2014
Event
1 November 2014
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The classroom note

Before generative AI tutors, researchers had already spent decades building intelligent tutoring systems — software that models what a specific student does and does not understand, and adapts accordingly — and, by 2014, enough separate evaluations existed to combine into a single statistical picture. A team led by Wenting Ma published that combination in the Journal of Educational Psychology in November 2014, pooling, as the paper's own abstract states, '107 effect sizes involving 14,321 participants' drawn from 'a search of major bibliographic databases.'

What the evidence says

The meta-analysis found intelligent tutoring systems produced 'greater achievement in comparison with teacher-led, large-group instruction' (an effect size of 0.42), against 'non-ITS computer-based instruction' (0.57), and against 'textbooks or workbooks' (0.35), but found 'no significant difference between learning from ITS and learning from individualized human tutoring' (-0.11) or small-group instruction (0.05) — a specific finding that these systems matched a private tutor rather than beating one. A second, independent meta-analysis, led by James Kulik and published in the Review of Educational Research in March 2016, pooled a smaller, more tightly filtered set of '50 controlled evaluations,' finding a larger median effect, '0.66 standard deviations... from the 50th to the 75th percentile,' while explicitly noting that the size of that effect 'depended to a great extent on whether improvement was measured on locally developed or standardized tests' — evaluations using researchers' own tests tended to show bigger gains than those using an external, standardised measure.

The implementation question

Both reviews are pooling evaluations of rule-based and model-tracing tutoring systems built for narrow domains — algebra, physics problem sets, reading comprehension — over years of development by specialist teams. Neither reviewed a generative, large-language-model tutor, which did not exist in this form at either publication date; a school reading these effect sizes as evidence for a 2026 AI chatbot is applying results measured on a different category of software, however similar the marketing language.

What holds and what fails

What holds, and is worth carrying forward, is Kulik and Fletcher's specific warning about measurement: an evaluation using its own test will tend to show a bigger effect than one using an independent, standardised test, a caution that applies to any tutoring product's self-reported results, generative or not. What fails is extending either meta-analysis's pooled effect size to a different technology; both reviews describe systems built and validated over long development cycles for narrow subjects, and neither offers evidence about open-ended generative tutoring.

  • Was the outcome measured with the vendor's own test or an independent, standardised one?
  • Is the tool being compared with a human tutor, a textbook, or ordinary large-group teaching — the comparisons differ hugely?
  • Is a pooled effect size from narrow, rule-based tutoring systems being applied to a generative AI product with no such evaluation?

Two large, independent meta-analyses agree that structured intelligent tutoring systems can match individual human tutoring on narrow academic tasks; neither says anything, on its own terms, about the open-ended generative tutors now being sold under the same name.

Sources & reading trail

Intelligent tutoring systems and learning outcomes: A meta-analysis ↗

Gives the pooled sample (107 effect sizes, 14,321 participants) and effect sizes of ITS versus teacher-led, computer-based, textbook, human-tutoring and small-group comparisons.

Source published: 1 November 2014 · Retrieved: 16 September 2026

Effectiveness of Intelligent Tutoring Systems ↗

Gives the 50-evaluation pooled median effect (0.66 SD) and the finding that effect size depended on whether local or standardised tests were used.

Source published: 1 March 2016 · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.