Listening: Tracking Multiple Speakers
8 min read
Sections 1 and 3 both involve more than one speaker, and a specific, recoverable type of error shows up in exactly these sections: hearing a piece of information correctly, but attributing it to the wrong speaker, which produces a wrong answer even though your comprehension of the actual words was accurate. This matters most for questions that specifically ask whose opinion or statement something was — a Matching question asking which of three students holds a particular view is testing attribution as much as comprehension, and getting the content right while assigning it to the wrong person still scores zero.
Before the audio for a multi-speaker section begins, use the available context — the question wording, any names given, the general setup described in instructions — to note how many speakers are involved and what role each one seems to play. A Section 1 conversation is usually two speakers in a clearly asymmetric relationship (a customer and a staff member, an applicant and an official), which makes attribution relatively easy since one person is usually asking and the other answering. Section 3 often involves two or three speakers in a more symmetrical relationship — two or three students discussing an assignment, sometimes with a tutor — where anticipating this structure in advance (three students, for instance, meaning three simultaneous or near-simultaneous viewpoints may be in play) primes you to track multiple voices from the first line rather than realising partway through that more than two people are speaking.
Voice alone is an unreliable way to track speakers, especially between speakers of similar age, gender, or accent, where two voices can sound genuinely difficult to distinguish for anyone who isn't a native listener of that accent. Content and conversational role are more reliable cues: in most exchanges, one person is asking a question and another is answering it, one person is proposing something and another is responding to the proposal, or one person's turn is clearly shorter and more reactive ("right," "I see," "that makes sense") than the other's. Tracking these conversational roles, rather than trying to memorise what each voice specifically sounds like, holds up better across a longer exchange.
In academic discussions particularly, speakers frequently propose a view, have it challenged, partly concede a point, and arrive at some kind of resolution — and the actual answer to a question often depends on where the exchange ends up, not on the first opinion stated. If Speaker A says "I think we should focus the presentation on the historical background," and Speaker B responds "That's useful context, but given our time limit, maybe we should focus more on the practical applications instead," and Speaker A then says "Yeah, that's fair, let's go with that," the exchange's actual conclusion — a focus on practical applications, agreed by both — is the answer a question about their final decision is looking for, not the historical-background suggestion raised first and later abandoned.
Ready to put this into practice?
Get 2 free scored Writing submissions, plus full Reading and Listening practice — no card required. Speaking is a separate credit-pack add-on.
Get 2 free scores