RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The atlas · 100 retrospective records ↗
School AI Atlas

The atlas / Assessment & integrity

Assessment & integrity / From the archive · April 2005 event · prepared 16 September 2026

ETS graded essays by algorithm years before chatbots existed

The testing body's own 2005 paper and current pages describe e-rater's role beside a human reader, not instead of one.

ets.orgprimary record

E-rater as a Quality Control on Human Scores (R&D Connections, No. 2)

Document
1 April 2005
Event
1 April 2005
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The classroom note

Long before a generative chatbot could write an essay on request, a testing organisation was already using software to grade one. In April 2005, ETS published E-rater as a Quality Control on Human Scores, a research note by William Monaghan and Brent Bridgeman describing why the organisation had invested in its e-rater automated essay evaluation system. For a student sitting the GRE's Analytical Writing section at that time, the practical fact was that a machine read the essay first, or alongside a person, rather than a single human reader deciding the score alone, a change driven by the logistics of large-scale testing rather than by any classroom decision.

What the evidence says

ETS's own 2005 account states the reasoning plainly: a single reader per essay 'does not produce reliable scores,' citing its own earlier research, so the organisation had long used at least two human readers per essay, at a cost the document calls 'substantial' once travel, training and compensation for a 'small army of educators' are added up. The paper reports that 'studies show a high level of agreement between the scores human raters assign to an essay and what e-rater awards,' citing a 2005 validation study of e-rater version 2.0, though the note itself does not restate that study's specific agreement statistic or sample size, only that it exists.

The implementation question

ETS's current pages, describing e-rater as retrieved on 16 September 2026, state that the engine still works 'in tandem' with human raters on high-stakes tests such as the GRE and TOEFL iBT, used on 'both the Issue and Argument prompts' and on 'Independent and Integrated Writing prompts' respectively, while a separate low-stakes product, Criterion, gives students automated feedback without a human check. The 2005 paper is candid about a limit that still applies: e-rater 'doesn't have the ability to read,' so it scores by pattern-matching against thousands of previously human-scored essays rather than by comprehension, which is why ETS has kept a human reader in the loop for its highest-stakes tests rather than replacing that reader outright.

What holds and what fails

What holds, across two decades of ETS's own description, is the hybrid design itself: automated scoring paired with human review, rather than automated scoring as a full replacement, for any test where the stakes are judged high enough to warrant it. What is likely to fail, editorially, is treating today's chatbot-detection debate as historically new; the 2005 paper already raised, and did not fully resolve, a concern about whether a system trained mostly on native-English-speaker writing would fairly score English-language learners, a question ETS's own document poses without reporting a settled answer.

  • Is an automated score here being used with a human check, as ETS's own high-stakes model does, or in its place?
  • What population was the scoring system trained on, and does it resemble the students it will now score?
  • Where a low-stakes product gives feedback without a human reader, does anyone treat that feedback as a grade anyway?

The debate over letting software grade writing did not start with generative AI; it is at least two decades old, documented in the testing industry's own research, and it settled, for its highest-stakes uses, on partnership with a human reader rather than replacement of one.

Sources & reading trail

E-rater as a Quality Control on Human Scores (R&D Connections, No. 2) ↗

ETS's own historical rationale for automated essay scoring, its cost logic, and reported agreement with human raters.

Source published: 1 April 2005 · Retrieved: 16 September 2026

e-rater Scoring Engine ↗

Current description of e-rater's use alongside human raters on the GRE and TOEFL iBT, and in the Criterion service, as retrieved.

Source published: Not established · Retrieved: 16 September 2026

Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.