7.3 Interactive Speaking: Real-Time Conversation Under a 35-Second Clock

Key Takeaways

  • Interactive Speaking replaced Listen, Then Speak on 1 July 2025; Listen, Then Speak and Read Aloud are both gone from the current test.
  • You hold a simulated conversation with an animated character and answer a series of 6-8 questions, with 35 seconds to record each answer.
  • Duolingo states that the questions are selected based on what you said in your previous answers, so the conversation genuinely follows your responses.
  • The official duration is 3:30 to 4:40 depending on how many questions are asked, and there is no preparation window before any of them.
  • Because there is no scratch paper and no prep time, answers must be built live: state a direct answer first, then add one reason or example.
Last updated: September 2026

7.3 Interactive Speaking: Real-Time Conversation Under a 35-Second Clock

Quick Summary: Interactive Speaking asks you to hold a simulated conversation with an animated character. Duolingo's official page states that you are asked a series of 6–8 questions, that you listen to the questions one at a time, and that you have 35 seconds to record yourself speaking the answer to each. The official duration for the whole set is 3:30 to 4:40, depending on how many questions come up. There is no preparation window, no visible text prompt, and no scratch paper. This section shows how to build a complete 35-second answer live.


What Replaced What, and When

Two changes took effect on 1 July 2025, and both are still missing from a great deal of DET prep material. Duolingo's own update article is unambiguous:

"Beginning July 1, 2025, the Duolingo English Test is adding Interactive Speaking. For Interactive Speaking, you will be asked a series of 6-8 questions in a simulated conversation with an animated character. You will listen to the questions one at a time and have 35 seconds to record yourself speaking the answer to each. The questions are selected based on what you said in your answers to previous questions. Interactive Speaking will contribute to your overall Duolingo English Test score and the speaking subscore. Interactive Speaking will replace the Listen, Then Speak question type. Additionally, the Read Aloud question type is being removed from the test to maintain the current test length."

TaskStatus on the current test
Listen, Then SpeakRemoved. Replaced by Interactive Speaking on 1 July 2025.
Read AloudRemoved on 1 July 2025 to keep the overall test length constant.
Interactive SpeakingCurrent. 1 set of 6–8 questions, 35 seconds each.

If a practice platform still serves you Listen, Then Speak or Read Aloud items, it is simulating a version of the test that no longer exists. Practising them is not harmful — Listen, Then Speak trained audio-only comprehension, and Read Aloud trained pronunciation — but neither will appear on test day, and neither trains the specific skill Interactive Speaking demands.


The Four Speaking Tasks, Side by Side

Only four question types feed the Speaking subscore, and they make very different demands. Getting their differences straight is the fastest way to stop practising the wrong pacing:

+---------------------+----------+-----------+-------------+------------------+
| Task                | Prompt   | Prep time | Talk time   | Count            |
+---------------------+----------+-----------+-------------+------------------+
| Speak About Photo   | Image    | 20 s      | up to 1:30  | 1 item           |
| Read, Then Speak    | Written  | 20 s      | up to 1:30  | 1 item           |
| Interactive Speaking| AUDIO    | NONE      | 35 s each   | 1 set of 6-8     |
| Speaking Sample     | Written  | 30 s      | up to 3:00  | 1 item           |
+---------------------+----------+-----------+-------------+------------------+

Read down the "prep time" column. Three of the four tasks give you a planning window; Interactive Speaking gives you none. Read down "talk time" and the contrast sharpens: everywhere else the challenge is sustaining speech; here it is compressing it. A 90-second answer structure delivered into a 35-second window gets cut off in the middle of its example.


The 35-Second Answer: A Structure That Actually Fits

Thirty-five seconds is roughly 70–85 words at a natural conversational pace. That is about four sentences. Any framework that needs an introduction, two developed points, and a conclusion will not fit — which is precisely why candidates who have drilled the PREP structure for Read, Then Speak often stumble here.

Use a compressed three-move structure instead:

[ 0-5 s  ] ANSWER    -> Answer the actual question in one direct sentence.
[ 5-25 s ] BECAUSE   -> One reason or one concrete example. Not both, not two of either.
[25-35 s ] ROUND OFF -> One short closing sentence. Stop talking.

Why "Answer First" Matters More Here Than Anywhere Else

Duolingo states that the questions are selected based on what you said in your answers to previous questions. The conversation is genuinely following you. That has a practical consequence most candidates miss: a vague opening does not just weaken one answer, it shapes what you are asked next.

  • Answer "Well, it depends on many factors, and I think different people have different opinions" and you have used eight seconds saying nothing, and given the next question nothing to build on.
  • Answer "I prefer studying in the library, mainly because my flat is too noisy" and you have used five seconds, established a position, and planted two follow-up hooks (the library, the noise) that you can develop when the next question arrives.

Worked Answer

Question heard: "Do you prefer studying alone or with other people?"

Answer (≈33 seconds, 74 words): "I definitely prefer studying alone, at least for anything that involves reading. When I study with friends, we usually end up talking about something else within about twenty minutes, so I lose the thread of whatever I was working on. Studying alone in the library, I can get through a full chapter without stopping. I do join a study group before exams, though, because explaining things out loud helps me notice what I have not really understood."

Notice the moves: a direct answer in the first clause, one specific reason with a concrete detail, and a short qualification that closes the answer naturally rather than trailing off. Notice also what it does not do — it does not open with "That's an interesting question," it does not list three reasons, and it does not run out of clock mid-sentence.


No Preparation Means No Notes: Building Answers in Working Memory

The DET forbids note-taking outright. Duolingo's published rules state "don't take notes or record questions" and "remove paper and notes from the testing space." Combined with zero preparation time, that means every Interactive Speaking answer is assembled in working memory while the clock runs.

The technique that makes this survivable is to commit, in advance, to a single-anchor habit:

[Question ends] --> Identify ONE anchor word from the question
                    |
                    v
        Anchor = the noun or verb the question is actually about
                    |
                    v
        Sentence 1: restate the anchor inside a direct answer
                    |
                    v
        Sentence 2-3: one reason or example attached to that anchor

If the question is "How has the way people shop changed in your country?", the anchor is shop. Sentence one: "Shopping has moved online almost completely for people my age." You are now committed, on topic, and four seconds in, without having planned anything.

[!TIP] Practise the first five seconds separately from everything else. Most Interactive Speaking damage happens before you have said a full sentence: silence, "umm", or a throat-clearing preamble. If your opening sentence is automatic, the remaining thirty seconds take care of themselves.


What Automated Speech Scoring Actually Measures

Duolingo describes how speaking responses are evaluated: models apply automatic transcription, natural language processing, and speech processing to assess content (relevance and development), coherence (organisation and ease of following), vocabulary and grammar (range and appropriateness), and fluency and pronunciation (whether you speak clearly, naturally, and understandably). It adds that these models are built around constructs such as intelligibility, syntactic range, and discourse coherence.

Translated into behaviour you can control inside 35 seconds:

+-----------------------+-----------------------+-------------------------+
| What is evaluated     | Do this                | Avoid this             |
+-----------------------+-----------------------+-------------------------+
| Content               | Answer the question   | Generic filler that    |
|                       | directly, add one     | could answer any       |
|                       | specific detail       | question               |
+-----------------------+-----------------------+-------------------------+
| Coherence             | One idea per sentence,| Restarting sentences;   |
|                       | clear connectors      | abandoning clauses      |
+-----------------------+-----------------------+-------------------------+
| Vocabulary & grammar  | Precise nouns and     | Repeating the same      |
|                       | verbs; one subordinate| three adjectives        |
|                       | clause                |                         |
+-----------------------+-----------------------+-------------------------+
| Fluency               | Pause at clause       | "Um", "uh", "like";     |
|                       | boundaries only       | mid-phrase hesitation   |
+-----------------------+-----------------------+-------------------------+
| Pronunciation         | Finish word endings;  | Mumbling; trailing off  |
|                       | speak at normal volume| at the end of sentences |
+-----------------------+-----------------------+-------------------------+

Pause Placement: The Fluency Detail That Is Actually Actionable

Duolingo's published fluency criteria include chunking — whether you pause naturally at the end of grammatical units — and breakdowns and repairs, meaning reliance on filler words, repetition, and starting and stopping. Some hesitation is explicitly described as natural; too much degrades the rating.

The distinction to internalise is where a pause falls, not whether one falls:

  • Natural: "I definitely prefer studying alone, [pause] at least for anything that involves reading." The pause sits at a clause boundary and sounds like thinking-in-sentences.
  • Disfluent: "I definitely prefer studying [pause] um [pause] alone." The pause sits inside a phrase, between a verb and its complement, and sounds like word-searching.

Both pauses are the same length. Only one reads as fluent.

Replacing Fillers Without Going Silent

Vocalised hesitations are the easiest thing to fix and the hardest habit to break. Substitute a short lexical hedge that buys the same half-second while still being language:

Instead ofSay
"Um... I think...""I'd say..."
"Uh, how do I explain this...""To put it simply..."
"Like, you know...""In other words..."
"...and, um, also...""On top of that..."

Recovery: When You Do Not Understand the Question

Because there is no text on screen and the question is spoken once in the flow of conversation, misunderstanding is a real possibility — especially on question four or five, when concentration is fraying. Silence is the worst available response: it produces no language for the engine to evaluate at all.

[Step 1] Salvage the anchors --> What content words DID you catch?
[Step 2] Generalise upward   --> What broad topic do those words belong to?
[Step 3] Answer the topic    --> Give a coherent 35-second answer on that topic.
[Step 4] Keep speaking       --> Never announce that you did not understand.

Worked recovery. Suppose the question is "Do you think municipal authorities should subsidise public transport fares?" and you catch only "public transport" and "should."

"I think public transport should definitely be a priority for cities. Where I live, the buses are reliable but quite expensive, so a lot of people drive even for short journeys. If the city made travel cheaper, or at least more predictable, I think far more people would leave their cars at home. It would reduce traffic and it would probably help with air quality too."

That answer is relevant, developed, coherent, and grammatically varied. It does not use the word subsidise once, and it does not need to.

[!CAUTION] There is a line between recovery and evasion. Duolingo states that off-topic responses are flagged by AI and reviewed by proctors, and that all spoken responses are reviewed for good-faith effort. Bridging from "public transport" to a coherent answer about public transport is good-faith recovery. Reciting a memorised paragraph about an unrelated topic because you did not understand is not, and it will read as off-topic.


Execution Checklist

  • Expect 6–8 questions at 35 seconds each, with a total set duration of 3:30 to 4:40.
  • Expect no preparation window. The recording clock starts when the question ends.
  • Answer in the first sentence. No preamble, no "that's a good question".
  • One reason or one example — not both. Thirty-five seconds holds about four sentences.
  • Plant hooks. The next question is selected from what you just said, so give it something specific to pick up.
  • Pause only at clause boundaries, and replace "um" with a short lexical hedge.
  • Never go silent. If you missed the question, answer the broad topic you did catch.
Test Your Knowledge

Which statement accurately describes Interactive Speaking as it currently appears on the Duolingo English Test?

A
B
C
D
Test Your Knowledge

A candidate is preparing from a study guide that includes drills for Listen, Then Speak and Read Aloud. What is the current status of those two question types?

A
B
C
D
Test Your Knowledge

During Interactive Speaking a candidate catches only the words 'public transport' from a question they otherwise did not understand. Which response best fits what Duolingo says it evaluates?

A
B
C
D