E-rater as a Quality Control on Human Scores (R&D Connections, No. 2)
- Document
- 1 April 2005
- Event
- 1 April 2005
- Retrieved
- 16 September 2026
The classroom note
Long before a generative chatbot could write an essay on request, a testing organisation was already using software to grade one. In April 2005, ETS published E-rater as a Quality Control on Human Scores, a research note by William Monaghan and Brent Bridgeman describing why the organisation had invested in its e-rater automated essay evaluation system. For a student sitting the GRE's Analytical Writing section at that time, the practical fact was that a machine read the essay first, or alongside a person, rather than a single human reader deciding the score alone, a change driven by the logistics of large-scale testing rather than by any classroom decision.
What the evidence says
ETS's own 2005 account states the reasoning plainly: a single reader per essay 'does not produce reliable scores,' citing its own earlier research, so the organisation had long used at least two human readers per essay, at a cost the document calls 'substantial' once travel, training and compensation for a 'small army of educators' are added up. The paper reports that 'studies show a high level of agreement between the scores human raters assign to an essay and what e-rater awards,' citing a 2005 validation study of e-rater version 2.0, though the note itself does not restate that study's specific agreement statistic or sample size, only that it exists.
The implementation question
ETS's current pages, describing e-rater as retrieved on 16 September 2026, state that the engine still works 'in tandem' with human raters on high-stakes tests such as the GRE and TOEFL iBT, used on 'both the Issue and Argument prompts' and on 'Independent and Integrated Writing prompts' respectively, while a separate low-stakes product, Criterion, gives students automated feedback without a human check. The 2005 paper is candid about a limit that still applies: e-rater 'doesn't have the ability to read,' so it scores by pattern-matching against thousands of previously human-scored essays rather than by comprehension, which is why ETS has kept a human reader in the loop for its highest-stakes tests rather than replacing that reader outright.
What holds and what fails
What holds, across two decades of ETS's own description, is the hybrid design itself: automated scoring paired with human review, rather than automated scoring as a full replacement, for any test where the stakes are judged high enough to warrant it. What is likely to fail, editorially, is treating today's chatbot-detection debate as historically new; the 2005 paper already raised, and did not fully resolve, a concern about whether a system trained mostly on native-English-speaker writing would fairly score English-language learners, a question ETS's own document poses without reporting a settled answer.
- Is an automated score here being used with a human check, as ETS's own high-stakes model does, or in its place?
- What population was the scoring system trained on, and does it resemble the students it will now score?
- Where a low-stakes product gives feedback without a human reader, does anyone treat that feedback as a grade anyway?
The debate over letting software grade writing did not start with generative AI; it is at least two decades old, documented in the testing industry's own research, and it settled, for its highest-stakes uses, on partnership with a human reader rather than replacement of one.
Sources & reading trail
ETS's own historical rationale for automated essay scoring, its cost logic, and reported agreement with human raters.
Source published: 1 April 2005 · Retrieved: 16 September 2026
Current description of e-rater's use alongside human raters on the GRE and TOEFL iBT, and in the Criterion service, as retrieved.
Source published: Not established · Retrieved: 16 September 2026
Departments, studies and vendor documents establish the record; the implementation reading and the boundary are School AI Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.