9.4 Tasks 3 & 4: Describing a Scene & Making Predictions

Key Takeaways

  • Tasks 3 and 4 use the same picture: Task 3 asks what is happening now, and Task 4 asks what will probably happen next (30 seconds preparation, 60 seconds speaking each).

  • CELPIP's feedback on its Level 8 sample says to address the listener directly and give an overview of the scene before describing details.

  • A consistent scanning route (foreground to background, or left to right) with clear location phrases helps the listener picture the scene.

  • Task 3 relies mainly on the present continuous for actions; Task 4 needs future forms and modals such as 'is about to', 'will probably', and 'might'.

  • Strong predictions link a visible clue to a likely result, for example: 'Because his shoelace is untied, he is about to trip.'

Last updated: October 2026

9.4 Tasks 3 & 4: Describing a Scene & Making Predictions

Quick Answer: CELPIP Speaking Tasks 3 and 4 operate as an interconnected pair based on the exact same full-colour illustration. In Task 3 (30 seconds prep, 60 seconds recording), your objective is to describe what is currently happening in the scene to someone who cannot see it, utilizing systematic spatial navigation and consistent Present Continuous verb forms. In Task 4 (30 seconds prep, 60 seconds recording), you examine the same visual scene to predict what will happen next, grounding your speculations in observable visual clues using future and modal constructions ("is about to", "is likely to", "will probably"). Structuring your visual progression logically and maintaining grammatical precision across both tasks is essential for securing CLB 9+.

The paired architecture of Tasks 3 and 4 tests two complementary communicative competencies: real-time spatial description and cause-and-effect extrapolation. Because both tasks rely on the same image stimulus, the 30-second preparation phase of Task 3 actually provides your first opportunity to survey the entire visual ecosystem. Candidates who understand the relationship between these two tasks can double their cognitive efficiency by identifying descriptive focal points in Task 3 that directly set up predictive storylines for Task 4.


The Paired Task Dynamic & Illustration Framework

CELPIP visual stimuli are detailed, stylized, full-colour cartoon illustrations packed with simultaneous human interactions, ambient environmental details, subtle comedic moments, and impending mishaps. Typical settings include a bustling community farmers market, an airport departure lounge, a public library, an active construction site, a dog park, a municipal swimming pool, or a neighbourhood street fair.

+-------------------------------------------------------------------------+
|               TASKS 3 & 4 PAIRED ARCHITECTURE                           |
+-------------------------------------------------------------------------+
|  VISUAL STIMULUS: One shared picture, typically a busy scene with       |
|  several people doing different things at the same time.                |
+------------------------------------+------------------------------------+
|  TASK 3: DESCRIBING A SCENE        |  TASK 4: MAKING PREDICTIONS        |
+------------------------------------+------------------------------------+
| - Prep: 30s | Recording: 60s       | - Prep: 30s | Recording: 60s       |
| - Temporal Focus: THE PRESENT      | - Temporal Focus: THE FUTURE       |
|   ("What is happening right now?") |   ("What will happen next?")       |
| - Core Grammar: Present Continuous | - Core Grammar: Modal verbs,       |
|   (is/are + verb-ing)              |   future forms (is about to / will)|
| - Strategy: Spatial navigation     | - Strategy: Cause-and-effect links |
|   (foreground -> background)       |   tied directly to visual evidence |
+------------------------------------+------------------------------------+

Task 3: Spatial Sequencing & Systematic Organization

In Task 3, CELPIP's instruction is to describe what is happening in the picture as well as you can to a person who cannot see it. CELPIP's feedback on its official Level 8 sample response notes two things that would have improved it: describing the scene directly to the other person, and giving an overview of the scene before the details. If you jump randomly around the image—mentioning a bird in the top right, then an apple on the ground, then a boy on the left—the rater cannot construct a coherent mental map. This chaotic "scattergun" approach severely damages your Content and Coherence score.

The Three Proven Spatial Scanning Patterns

To project an organized, structured delivery, select one of three systematic scanning patterns during your 30-second preparation and stick to it strictly:

  1. The Depth Scan (Foreground to Background): Begin with the dominant figures in the immediate foreground, transition to ongoing interactions in the midground, and finish with setting details in the distant background.
  2. The Horizontal Sweep (Left to Right): Anchor the overall setting, then systematically describe the action cluster on the left, transition across the central figures, and conclude with the activity on the right.
  3. The Anchor & Cluster Scan: Establish the global context ("This vibrant illustration depicts a crowded community farmers market on a sunny morning"), analyze the central focal point, and then explore two secondary peripheral scenes.
                     THE SPATIAL NAVIGATION GRID

     [ Upper Left Background ]    [ Upper Centre / Sky ]    [ Upper Right Background ]
     - Banner hanging between     - Flying kites / birds     - Food truck ordering line
       two brick buildings        - Distant hills / clouds   - Park gazebo

     [ Midground Left ]           [ Midground Centre ]      [ Midground Right ]
     - Vegetable vendor weighing  - Cobblestone pathway     - Acoustic guitarist playing
       organic produce            - Cyclist dodging puddle   - Couple dancing

     [ Foreground Left ]          [ Foreground Centre ]     [ Foreground Right ]
     - Toddler reaching toward    - Man carrying groceries   - Open guitar case with
       towering fruit pyramid       with untied shoelace       stray dog sniffing food

Spatial Preposition & Directional Phrase Bank

Mastering precise locational language is essential for demonstrating lexical range:

  • Foreground: "In the immediate foreground...", "Right at the bottom centre of the image...", "Down in the lower left corner..."
  • Midground: "Just behind them along the central walkway...", "Over in the centre of the square...", "Positioned directly between the two stalls..."
  • Background: "In the far background towards the upper right...", "Up against the rear brick wall...", "In the distance, hovering above the tree line..."
  • Relative Position: "Directly adjacent to...", "Flanked on either side by...", "In close proximity to...", "Perched precariously atop..."

Task 3 Grammar: Present Continuous Mastery

The fundamental grammatical requirement of Task 3 is the Present Continuous tense (subject + auxiliary 'be' + present participle). Because you are describing active physical motions occurring in front of your eyes, all dynamic behaviours must be framed in this tense.

Critical Grammatical Distinctions in Task 3

Verb TypeIncorrect UsageCorrect CLB 9+ Usage
Dynamic Human Action"A young boy chases a runaway dog." (Simple Present ❌); "A boy chasing dog." (Dropped Auxiliary ❌)"A young boy is chasing a runaway golden retriever across the lawn." (Present Continuous ✅)
Simultaneous Dynamic Actions"Two men talk and drink coffee." (Flat Simple Present ❌)"Two elderly gentlemen are chatting animatedly while sipping coffee on a bench." (Coordinated Continuous ✅)
Stative Background Elements"A wooden cart is being near the tree." (Ungrammatical ❌)"There is a rustic wooden cart beside the oak tree", "A wooden cart is parked beside the tree", or "A cart is standing near the tree" (all ✅)

Caution

Dropping the auxiliary verb ("A vendor selling apples" instead of "A vendor is selling apples") is one of the most pervasive grammatical slips among intermediate candidates. Ensure every singular subject is paired with "is" and every plural subject with "are".


Task 4: Predictive Logic & Speculative Phrasing

When the system advances to Task 4, the exact same illustration remains on screen, but the prompt changes dramatically: you are now asked to predict what is going to happen next.

Candidates who fail to shift cognitive gears often make the fatal error of re-describing the picture in the present tense. Task 4 evaluates your ability to speculate, extrapolate, and hypothesize using future grammatical constructions.

The Visual Cause-and-Effect Formula

Every prediction in Task 4 must be anchored in observable visual evidence. High-scoring candidates use a two-step formula:

Visual Observation (Cause) + Speculative Modal Marker → Logical Future Outcome (Effect)

  • Weak (Unanchored guess): "I think a police car will arrive and arrest someone." (No evidence in the picture)
  • Strong (CLB 9+ Evidence-Based): "Notice that the man carrying two heavy grocery bags in the centre path has an untied shoelace dragging beneath his foot; because of this, he is almost certainly about to trip and send his groceries tumbling across the pavement."

The Speculative Grammar Spectrum

Demonstrate lexical and grammatical sophistication by modulating your degree of certainty:

+-------------------------------------------------------------------------+
|                   THE SPECULATIVE GRAMMAR SPECTRUM                      |
+-----------------------+-----------------------+-------------------------+
| CERTAINTY LEVEL       | MODAL / PHRASAL FORMS | AUTHENTIC SPOKEN USE    |
+-----------------------+-----------------------+-------------------------+
| Immediate / Highly    | - is about to...      | "The toddler is on the  |
| Probable Action       | - is on the verge of..| verge of toppling the   |
|                       | - will almost certainly| entire pyramid of fruit."|
+-----------------------+-----------------------+-------------------------+
| Likely / Probable     | - is likely to...     | "The stray dog will     |
| Consequence           | - will probably...    | probably snatch the     |
|                       | - appears poised to...| sandwich right out of   |
|                       | - looks as though...  | the open guitar case."  |
+-----------------------+-----------------------+-------------------------+
| Tentative / Possible  | - might inadvertently | "The two ladder workers |
| Hypothesis            | - could potentially.. | could potentially drop  |
|                       | - may end up -ing...  | the banner on shoppers."|
+-----------------------+-----------------------+-------------------------+

Comparative Scenario: Community Farmers Market

To see how Tasks 3 and 4 complement each other, examine how the same four visual focal points within a busy community farmers market illustration are handled across both tasks.

+-------------------------------------------------------------------------+
|  SCENE DESCRIPTION: A BUSTLING COMMUNITY FARMERS MARKET                 |
|  - Foreground Left: A fruit vendor weighing tomatoes while an unattended|
|    toddler reaches toward a precarious pyramid of stacked oranges.      |
|  - Centre Pathway: A shopper carrying two overflowing tote bags with an |
|    untied shoelace trailing directly under his sneaker.                 |
|  - Foreground Right: A street musician strumming a guitar with his open  |
|    case resting on the ground, where a stray beagle is sniffing food.   |
|  - Background Upper Right: Two volunteers on wobbly stepladders trying  |
|    to hang a large canvas welcome banner between two light posts.       |
+-------------------------------------------------------------------------+

Model Spoken Transcript: Task 3 — Describing the Scene (Written to Target CELPIP 10; Coaching Estimate)

[0:00 – 0:12] "This lively illustration depicts a bustling outdoor farmers market on a bright weekend morning, packed with shoppers and local vendors.

[0:12 – 0:28] Starting in the foreground on the left-hand side, a friendly produce vendor wearing a striped apron is carefully weighing heirloom tomatoes on a hanging scale. Right beside his stall, an unattended toddler in yellow overalls is reaching his hands up toward a towering, precarious pyramid of oranges.

[0:28 – 0:44] Moving over to the centre walkway, a busy customer is power-walking through the crowd while carrying two overflowing grocery totes filled with fresh baguettes and greens. Down on the cobblestones, his left shoelace is completely untied and trailing behind his sneaker.

[0:44 – 0:58] Meanwhile, on the right side of the path, a street musician is strumming an acoustic guitar while a couple is swaying to the music. On the pavement in front of him, an open guitar case contains his tip jar and an unwrapped sandwich, which an inquisitive beagle is currently sniffing. Finally, up in the background, two volunteers are balancing on stepladders trying to secure a massive canvas welcome banner."

Model Spoken Transcript: Task 4 — Making Predictions (Written to Target CELPIP 10; Coaching Estimate)

[0:00 – 0:14] "Based on the precarious actions unfolding across this farmers market scene, several dramatic incidents are bound to happen within the next few moments.

[0:14 – 0:29] First, look at the toddler on the lower left who is tugging at the bottom of the fruit stand. In all likelihood, he is about to pull one of the foundational oranges loose, which will inevitably cause the entire citrus pyramid to collapse and roll across the cobblestones, catching the busy vendor completely off guard.

[0:29 – 0:44] Simultaneously, out on the main pathway, because the hurried shopper is power-walking through the crowd without watching his footing, he will almost certainly step on his untied shoelace. This will cause him to trip headfirst, sending his baguettes, produce, and groceries flying into the crowd.

[0:44 – 0:58] Over by the musician, that opportunistic beagle appears poised to snatch the unattended sandwich right out of the guitar case before the guitarist even notices. Lastly, up on the stepladders, the volunteer on the left looks as though he might lose his balance, which could force them to drop the banner onto the shoppers passing below."


Coaching Analysis Against CELPIP's Four Categories

Assessment DimensionTask 3 Evaluation (Describing Scene)Task 4 Evaluation (Making Predictions)
Content & CoherenceEvaluates whether the candidate followed a structured spatial trajectory (foreground-left → centre → foreground-right → background) rather than jumping randomly.Evaluates whether each prediction is logically derived from visual clues rather than arbitrary, ungrounded speculation.
VocabularyRewards rich descriptive and spatial vocabulary ("heirloom tomatoes", "untied shoelace", "inquisitive beagle", "precarious pyramid").Rewards predictive and speculative phrasing ("bound to happen", "in all likelihood", "appears poised to", "inevitably cause").
ListenabilityPrioritizes smooth, continuous speech flow and clear articulation of consonant clusters and participles ("weighing", "strumming", "balancing").Prioritizes natural cadence and sentence stress when expressing degrees of certainty and probabilistic emphasis.
Task FulfillmentDescribes the scene to the listener, starts with an overview, and uses most of the 60 seconds.Answers the actual question (what will probably happen next) with several distinct, evidence-based predictions.
Test Your Knowledge

What is the most effective organizational strategy when describing an illustration in Speaking Task 3?

A

Follow one consistent route—foreground to background or left to right—using clear location phrases.

B

Jump rapidly between the most colourful objects across different corners of the image without mentioning their locations.

C

Focus exclusively on a single person's facial expression for the entire 60 seconds.

D

Describe what the illustrator might have eaten for breakfast before drawing the picture.

Test Your Knowledge

How should a test taker formulate high-scoring predictions in Speaking Task 4?

A

Invent dramatic fictional disasters that have no connection to the visual details shown in the image.

B

Repeat the exact present continuous descriptions spoken in Task 3 without altering the verb tenses.

C

Link each prediction to a visible clue in the picture, using forms such as 'is likely to' or 'is about to'.

D

Declare with absolute 100% certainty that only one single event will occur in the scene over the next twenty years.

Test Your Knowledge

Which tense is most appropriate for describing the actions people and animals are performing in a CELPIP Task 3 picture?

A

Past Perfect ('A vendor had sold apples')

B

Present Continuous ('A vendor is arranging fresh apples')

C

Future Continuous ('A vendor will be arranging apples')

D

Simple Past ('A vendor arranged fresh apples')

Sections you finish are checked off in the contents.