IELTS Format: How Speaking Is Actually Scored (Interview Structure)
6 min read
A specific misconception trips up a lot of candidates preparing for Speaking: the assumption that the examiner is grading you the way a written exam is graded, checking off correct and incorrect answers against a fixed key, or that there's a hidden "target vocabulary" the examiner is listening for you to use. Neither is true. Speaking is scored live by the same person conducting your interview, assessing your actual spoken performance in real time against four defined criteria, with no separate scoring examiner listening in from another room and no fixed script of "correct" responses to compare you against — there genuinely is no right answer to most Speaking questions, only a range of quality in how any given answer is delivered.
The interview itself follows a structured three-part format, and understanding what's actually being assessed in each part changes how you should approach it. Part 1 involves short, familiar questions about yourself — home, work or study, hobbies, familiar topics — lasting several minutes, and primarily gives the examiner an initial read on your basic fluency and range before the more demanding parts. Part 2 is the individual long turn: you're given a task card with a topic, one minute to prepare notes, and then are expected to speak for one to two minutes largely uninterrupted, which is the part of the interview most directly testing your ability to organise and sustain extended speech without the scaffolding of back-and-forth questions. Part 3 returns to two-way discussion, but shifts to more abstract, analytical questions connected to the Part 2 topic, testing your ability to discuss ideas, justify opinions, and handle more complex, less personal content than Part 1.
The examiner works from a structured format with prescribed timing for each part, but has genuine flexibility in the exact follow-up questions used within that structure — this is why two candidates who both do a Part 2 about a "memorable trip" will still get somewhat different specific questions in Part 1 and Part 3, even though the overall shape of the interview is standardised. This flexibility exists specifically so the examiner can follow up naturally on what you actually say, rather than reading from a completely fixed script regardless of your answers — which also means genuinely engaging with the conversation, rather than delivering a memorised, generic answer that ignores what was actually asked, tends to produce a more natural, better-assessed performance.
Scoring happens progressively across the entire interview rather than being decided only at the end or based on a single strong or weak moment. Throughout all three parts, the examiner is continuously forming an impression of your performance against Fluency and Coherence (how smoothly and logically you speak, without excessive hesitation or a disorganised structure), Lexical Resource (the range and precision of your vocabulary, including how well you handle topics you may not have specific vocabulary prepared for), Grammatical Range and Accuracy (the variety of grammatical structures you use and how accurately you control them), and Pronunciation (clarity, stress, intonation, and the overall ease with which a listener can follow you). Your final band for Speaking reflects the overall pattern across all three parts and all four criteria, not simply your best moment or your average calculated from a couple of standout sentences.
Ready to put this into practice?
Get 2 free scored Writing submissions, plus full Reading and Listening practice — no card required. Speaking is a separate credit-pack add-on.
Get 2 free scores