IELTS Format: How Examiners Are Trained (Scoring Consistency)
6 min read
A recurring worry among candidates, especially after an unexpectedly low Writing or Speaking score, is that the result mostly reflects which particular examiner happened to mark or interview them, rather than a consistent, reproducible standard. This concern is understandable — Writing and Speaking are human-scored, unlike the purely mechanical scoring of Reading and Listening, and any human-judged assessment naturally raises questions about consistency. It's worth understanding, though, that IELTS has built a fairly extensive system specifically to address this concern, rather than leaving examiner judgment entirely unchecked.
Before being approved to examine at all, prospective IELTS examiners go through standardised initial training and a certification process, using the same detailed, publicly available band descriptors and the same training materials regardless of which country or test centre they'll eventually work at. This matters because it means an examiner training in one country is being calibrated against literally the same written standard, with the same benchmark sample scripts and recordings, as an examiner training in a completely different country — the intention is that a candidate's script or interview should be assessed against the same fixed criteria no matter where in the world it happens to be marked.
Certification isn't a one-time event after which an examiner is simply left to mark independently forever. Examiners are periodically re-certified, and their live examining and marking is monitored on an ongoing basis against current benchmark standards, specifically to catch any gradual drift — a tendency, for instance, to mark slightly more leniently or strictly over time than the fixed descriptors intend — before it becomes a larger, systematic problem. This ongoing monitoring, rather than a single initial qualification, is what's meant to keep examiner standards aligned with each other over years of active examining, not just at the point of certification.
For Speaking specifically, interviews are typically recorded, and a proportion of these recordings are reviewed separately from the live scoring decision as part of ongoing quality assurance. This creates a genuine, independent check on an individual examiner's consistency beyond the live, in-the-moment judgment made during the interview itself — a second reviewer working from the same recording and the same band descriptors provides a way to detect and address any individual examiner whose live scoring pattern has started to diverge from the standard.
For Writing, scripts are marked against detailed, criterion-specific band descriptors rather than an examiner's holistic, general impression of the essay's overall quality — the intention is that two different examiners working from the same fixed descriptors, applied to the same script, should reach the same or a very closely matching band. To verify this in practice, examiner marking is also periodically cross-checked against other examiners' marking of the same or comparable benchmark scripts, which serves the same consistency-monitoring purpose for Writing that recorded-interview review serves for Speaking.
Ready to put this into practice?
Get 2 free scored Writing submissions, plus full Reading and Listening practice — no card required. Speaking is a separate credit-pack add-on.
Get 2 free scores