
The classroom note
On 31 January 2023, OpenAI published a free tool for judging whether text was AI-written, aimed at a concern already spreading through schools. The release post offered it "to get feedback on whether imperfect tools like this one are useful" and named "academic dishonesty" as one problem it might address. Fewer than six months later, the same page carries a one-line update: "As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy." A teacher who had started using it over the spring term lost access before the next school year began.
What the evidence says
OpenAI's release post states the classifier's measured performance directly: it "correctly identifies 26% of AI-written text (true positives) as 'likely AI-written,' while incorrectly labeling human-written text as AI-written 9% of the time (false positives)." That is a company reporting its own tool catching roughly a quarter of AI-written text in its own evaluation, on a "challenge set" of English texts whose size and source go undetailed. The post warns the classifier "should not be used as a primary decision-making tool," is "very unreliable on short texts," and performs "significantly worse" outside English. A separate, undated educator help page, current as retrieved 16 September 2026, goes further: OpenAI's own testing found that training a detector "labeled human-written text like Shakespeare and the Declaration of Independence as AI-generated," and that detectors could "disproportionately impact" students learning English as a second language.
The implementation question
The mechanism a school needs is not whether a detector exists but what a 26% true-positive rate means for a single accusation. A tool that misses roughly three-quarters of AI-written text in the vendor's own test, while mislabelling nine in a hundred human texts, is a poor basis for any single disciplinary decision, precisely OpenAI's own conclusion, stated on its help page: "Do AI detectors work? In short, not in our experience." This is the vendor withdrawing not just a product but the claim that detection at scale was achievable with the method it tried.
What holds and what fails
What holds is the dated, quoted sequence: a company published its own accuracy figures, then withdrew the tool five and a half months later citing that same low accuracy. What fails is any assumption that a similarly branded detector from another vendor has solved the problem OpenAI abandoned; these two pages describe only OpenAI's classifier, not the field generally.
- Does the school rely on any AI-text detector, and has its accuracy been checked against OpenAI's experience?
- Would the school's policy survive OpenAI's own finding about non-native English writers and formulaic prose?
- If OpenAI's own classifier caught roughly a quarter of AI text while flagging some human text, what false positive rate would the school find acceptable elsewhere?
A company naming its own detector's failure rate, then withdrawing the product within half a year, is one of the clearest documented cases of a vendor changing its own claim in public; the two pages together caution against any newer tool making a similar promise without similar numbers attached.
Sources & reading trail
OpenAI's own release post stating the classifier's 26% true-positive and 9% false-positive rates, plus an inline update recording withdrawal on 20 July 2023 for low accuracy.
Source published: 31 January 2023 · Retrieved: 16 September 2026
Living educator FAQ stating OpenAI's own conclusion that detectors do not reliably work and mislabelled human-written texts including Shakespeare during testing.
Source published: Not established · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.