4.4 Listening Task Formats, Recording Length & the Two-Playback Rule
Key Takeaways
- EF states that each test taker answers approximately 50 listening questions and that the number of listening prompts varies by session.
- EF SET listening recordings run between 20 seconds and 5 minutes, and each recording can be listened to twice.
- After a recording finishes, EF gives you as much time as you need to answer its questions — but that time and both playbacks all run on the same 25-minute section clock.
- EF's design specifications use three listening task models: a single speaker with multiple-choice or multiple-option questions plus completions, a dialogue with multiple-choice questions, and a speaker-match task.
- Because the section is adaptive, both the length of the recordings and the difficulty of the questions change according to how you are performing.
Listening Task Formats, Recording Length & the Two-Playback Rule
Quick Summary: The single most consequential fact about the EF SET listening section is one that most preparation material gets wrong: every recording can be listened to twice. EF states it plainly. Knowing it changes your entire approach — from a panicked one-shot transcription effort into a deliberate two-pass routine, budgeted against a 25-minute clock that both playbacks share.
The Published Numbers
| Published fact | Value | Consequence for your plan |
|---|---|---|
| Questions per test taker | Approximately 50 | About 30 seconds per question if spread flat across 1,500 seconds. |
| Number of recordings | Varies by session | The adaptive panel decides; there is no fixed count. |
| Recording length | 20 seconds to 5 minutes | A single long track played twice can consume a fifth of the section. |
| Playbacks per recording | Two | The replay is available on every recording, not just short ones. |
| Answering time | "As much time as he needs" after the recording | Untimed per task — but inside the 25-minute section limit. |
| Adaptivity | Recording length and question difficulty adapt | Longer, harder audio is a signal that you routed upward. |
[!IMPORTANT] The replay is not free. EF removes the per-task timer, not the section timer. Replaying a 4-minute lecture costs four minutes you will not get back. The skill being tested here is not endurance; it is judgement about when a second pass is worth its price.
The Three Listening Task Models
EF's academic and technical development report specifies three task-model families for listening.
1. Monologue With Multiple-Choice or Multiple-Option Questions
One speaker delivering a continuous stretch of speech: an announcement, a voicemail, a briefing, a mini-lecture. At higher levels EF's specifications add completion items alongside the multiple-choice set. This is the format that Section 5.1's real-time mapping technique is built for, and the format where the second playback earns its cost most often.
2. Dialogue With Multiple-Choice Questions
Two speakers in exchange. Items typically probe who wants what, what relationship holds between the speakers, and what a turn was doing rather than what it literally said — the pragmatic ground covered in Sections 6.1 and 6.2. Dialogues sit at the shorter end of the duration range, so they are usually resolvable on a single pass.
3. Speaker Match
Several short extracts from different speakers must be matched to a set of options — an opinion, a role, a situation, or a stated purpose. This format has its own failure mode: candidates lock in an early match, then discover a later speaker fits it better. Handle it as follows:
- Hold matches provisionally through the first pass; write nothing in stone.
- Note the single most distinctive content word or attitude marker for each speaker.
- Assign the confident matches first, then resolve the remainder by elimination.
- Use the second playback selectively, on the two extracts you could not separate.
Accents Are Part of the Construct
EF states that the listening section is "composed of a series of recorded texts read by American, British English, and Australian English speakers." This is not incidental variety; EF lists "number and accents of speakers" among the explicit levers it uses to make listening items harder.
That has a direct preparation consequence. Training exclusively on one variety — a single news broadcaster, a single podcast host — leaves you exposed on a third of the material. Chapter 4.2 covers the specific phonological contrasts (the trap–bath split, non-rhotic r, intervocalic flapping, Australian high rising terminal); the point here is structural: all three varieties are in scope on every sitting, and the engine will not warn you which one is coming.
Budgeting the Two Playbacks Across 25 Minutes
Here is the arithmetic that decides your section. Assume a mixed set of recordings:
Section budget: 1,500 seconds
Short dialogue (~30s) × 1 play = 30s + 3 items × 15s answering = 75s
Announcement (~45s) × 2 plays = 90s + 2 items × 15s answering = 120s
Extended monologue (~4 min) × 2 plays = 480s + 8 items × 20s answering = 640s
─────────────────────
three prompts ≈ 835 seconds
Two long monologues replayed in full will consume most of a section on their own. The discipline that follows:
| Recording type | Default playback plan | Replay only when |
|---|---|---|
| Short dialogue (20–45s) | One pass, answer immediately | A number, a name or a negation is genuinely unresolved |
| Announcement / voicemail (30–90s) | One pass for function, replay if any figure is at stake | Times, platform numbers, prices or dates are the item's target |
| Extended monologue (2–5 min) | Map on pass one, replay for specific unresolved items | Two or more items still hang on a detail you did not capture |
| Speaker match | One pass to hold provisional matches | Two extracts remain interchangeable after elimination |
[!CAUTION] The passive-second-listen trap. The most common way to lose a listening section is to replay everything by reflex, sitting through a second pass without a specific question in mind. A second playback should be a targeted search, not a re-experience. Before you press play again, name the thing you are listening for.
Pre-Audio Preparation and the Sound Check
Before the listening section starts, the interface takes you through a sound check. Use it properly:
- Play the calibrated sample and set your system and headphone volume then, not during the first scored item.
- Confirm both channels are audible — a dead earcup on a dialogue task is fatal.
- Modern browsers block unprompted audio autoplay, so the test requires an explicit click to initialise playback; expect that prompt rather than assuming the audio has failed.
Then, on each prompt, spend the short window before playback reading the item stems. That five-to-ten-second preview is what converts the first playback from passive listening into a targeted search, and it is the cheapest score improvement available in the whole section.
[!TIP] A rule of thumb worth memorising: preview the stems, listen once for structure, decide whether a replay would answer a named question, and only then use it. Candidates who replay by default finish the section; candidates who replay on purpose finish it with time in hand.
How many times may a test taker listen to each EF SET listening recording, and how long do the recordings run?
A four-minute academic monologue carries eight questions, and you have resolved six of them after the first playback. What is the best use of the second playback?
Which set of task models does EF's design documentation specify for the listening section?
Which English varieties appear in EF SET listening recordings, and what does EF say about their role in item difficulty?