RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Assessment & integrity

Assessment & integrity / Reference note · Reference note · prepared 16 September 2026

A detector's accuracy claim meets a documented bias finding

A vendor's own false-positive claim, an independent bias study and exam-board guidance together show why one score should not decide a case.

Visual published with the cited source for this record: A detector's accuracy claim meets a documented bias finding
Visual published with the cited source, shown for identification of the record. Credit: turnitin.com · source page ↗ Rights: owner-review-pending.

The classroom note

A teacher opens an AI-detection score next to a student's essay and has to decide what the number means before deciding anything about the student. Reading the detector vendor's own account of its accuracy alongside independent research on how these tools fail, as retrieved on 16 September 2026, turns that single score into a more careful question: accurate for whom, and under what conditions?

What the evidence says

Turnitin's own account of false positives, published 16 March 2023 by its Chief Product Officer, states the company has aimed for high accuracy alongside "a less than 1% false positive rate" - a vendor's claim about its own model, not an independent measurement. A separate preprint study, first posted in April 2023, tested several widely used GPT detectors on writing from native and non-native English speakers and found the tools "consistently misclassify" non-native English writing as AI-generated while accurately identifying native writing; simple prompt-rewriting, the authors add, can also let AI-generated text evade detection entirely. Neither source proves the other wrong: a low aggregate false-positive rate can still sit alongside a large, patterned error concentrated in one group of students.

The implementation question

The Joint Council for Qualifications' guidance, updated 30 April 2025, effectively reconciles the two findings by refusing to treat any single score as decisive. It states that detection tools "vary in accuracy" depending on the tool and version used, the proportion of AI content, and other factors including "an individual's English language competency," and recommends using "more than one detection tool" when misuse is suspected, with detection forming "part of a holistic approach" rather than standalone proof. That is a process requirement, not a technical fix: it shifts the decision from a number to a documented pattern of evidence a teacher can defend.

What holds and what fails

What holds: a detection score can legitimately prompt a closer look, especially alongside comparison with a student's earlier, supervised work. What fails: reading a percentage on a marketing page as a probability that applies equally to every student in a class, when the cited research ties error specifically to language background. This is an editorial checklist, not a malpractice ruling: schools remain bound by their own exam board's process, and no score alone should determine an outcome for a named student.

  • Is this score being treated as the sole evidence, or one input among several, including the student's own account?
  • Does the vendor's accuracy claim state the sample it was measured on, and does it address non-native English writers?
  • Has more than one detection tool been used, as the guidance recommends, before any action is taken?

A detector narrows attention; it does not settle a case. The research and the guidance both point the same way: toward more evidence, not a faster verdict.

Sources & reading trail

Understanding false positives within our AI writing detection capabilities ↗

Vendor blog states a target of a less than 1% false-positive rate for its AI-writing detector.

Source published: 16 March 2023 · Retrieved: 16 September 2026

GPT detectors are biased against non-native English writers ↗

Preprint finds several GPT detectors consistently misclassify non-native English writing as AI-generated, and that prompting can bypass detection.

Source published: 6 April 2023 · Retrieved: 16 September 2026

AI Use in Assessments: Your role in protecting the integrity of qualifications ↗

States detection-tool accuracy varies with factors including English language competency, and recommends using more than one tool within a holistic process; shown as updated 30 April 2025 at retrieval.

Source published: Not established · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.