Understanding Task 3
Task 3 presents you with a photograph — typically a scene from everyday Canadian life — and asks you to describe what you see. You have 30 seconds to prepare and 60 seconds to speak. The photograph will usually show people doing something: working, socializing, being outdoors, shopping, celebrating, or engaged in a community activity. Your job is to describe the scene in an organized, specific, and natural way — without guessing at things that are not clearly visible, and without simply listing what you see without any structure.
Task 3 and Task 4 share the same photograph. This means the preparation time you use in Task 3 also prepares you for Task 4. Smart candidates use their 30 seconds not only to plan their description but also to identify the main visible elements they will use in their Task 4 prediction. Planning both tasks during Task 3 preparation is a key efficiency habit that reduces preparation pressure in Task 4.
The Smart Scanning Order
Before speaking, spend your 30 preparation seconds scanning the photograph in a deliberate order. Random scanning leads to random, unorganized descriptions. The following scanning sequence produces organized, structured descriptions:
- Overall scene (2 seconds): What is the setting? Indoor or outdoor? What type of location is this? (park, office, kitchen, street, store, etc.)
- People (10 seconds): How many people? What are they doing? What approximate age group? What are they wearing? What are their expressions or body language?
- Objects and environment (8 seconds): What significant objects, furniture, equipment, or environmental features are visible? What is in the background?
- Overall impression (5 seconds): What is the mood or atmosphere of the scene? What does the scene suggest about the occasion or relationship between people?
The ODIC Description Structure
Use the ODIC structure to organize your 60-second Task 3 response: Overview, Details, Inference, Connection.
O — Overview (8–10 seconds): Begin with a broad, orienting sentence that describes the overall scene — where it is, what type of activity is happening, and approximately how many people are involved. This gives the listener a mental picture before you add detail.
D — Details (30–35 seconds): Describe three to four specific, concrete details from the photograph. Focus on actions (what people are doing, using present continuous tense: "A woman is talking on the phone"), objects (what is visible in the environment), and expressions or body language if visible. Prioritize what is most prominent and relevant in the photograph.
I — Inference (10–12 seconds): Make one logical, limited inference about what the scene suggests. This demonstrates higher-order thinking and adds a reflective dimension to your description. Crucially, signal your inference clearly so the rater knows you are interpreting, not describing: "This suggests that..." / "It looks as though..." / "The scene gives the impression that..."
C — Connection (5–8 seconds): End with a brief statement that connects the scene to a familiar context or your own experience. This optional step adds naturalness and coherence to your conclusion. "Scenes like this are common in..." / "I have been in a similar situation when..." / "It reminds me of..."
Grammar for Scene Description: Present Continuous and Passive Voice
Task 3 descriptions should primarily use present continuous tense for actions (what people are doing at the moment of the photograph) and simple present for states (what exists in the scene). Mixing in passive voice naturally shows grammatical range.
| Structure | Example |
|---|---|
| Present Continuous (actions) | "A woman is pouring coffee into a mug." |
| Present Continuous (states) | "Two men are standing near a window." |
| Simple Present (existence) | "There are three tables arranged in a row." |
| Passive Voice (focus on object) | "Several boxes are stacked against the wall." |
| Hedged Inference | "It appears that they are in the middle of a meeting." |
| Appearance Language | "The people look relaxed and seem to be enjoying themselves." |
Vocabulary for Describing Scenes
Location and Setting Words
- in the foreground / background / left / right / centre of the image
- in what appears to be a [kitchen / office / park / café / classroom]
- outdoors in a natural setting / indoors in a professional environment
- the space looks [spacious / cozy / busy / calm / well-organized]
People Description Words
- a group of / a couple of / several / what appears to be [number] people
- they appear to be [young adults / middle-aged / older / children]
- dressed in [casual / professional / outdoor / work] clothing
- their body language suggests [concentration / enjoyment / collaboration / urgency]
Action Description Words
- is engaged in / is involved in / appears to be [activity]
- is interacting with / is focused on / is directing attention toward
- is gesturing toward / is pointing at / is holding / is reaching for
The Most Common Task 3 Mistakes
- Guessing at invisible details: Only describe what is clearly visible. If you cannot see someone's face clearly, do not describe their emotion — say "they appear engaged" based on body language. If you cannot read text in the photo, do not invent what it says. Inaccurate descriptions hurt Task Fulfillment.
- Listing without structure: Responses that simply list what is visible ("I see a table and chairs and people and some food and a window") have very poor Coherence scores. Always use the ODIC structure.
- Using only simple sentences: "There is a woman. She is sitting. She has a laptop." All simple sentences indicate limited grammatical range. Combine ideas: "A woman is sitting at a desk with a laptop open in front of her, and she appears to be focused on her work."
- Stopping before 60 seconds: If you describe the scene in 35 seconds and stop, you have not fulfilled the task. Use the Inference and Connection sections to fill the remaining time with quality content.
Key Takeaways from Lesson 8
- Scan the photograph in a deliberate order: overall scene → people → objects → overall impression.
- Use the ODIC structure: Overview → Details → Inference → Connection.
- Use present continuous for actions, simple present for states, and passive voice for objects.
- Signal inferences clearly with hedging language: "It appears that..." / "This suggests..."
- Never guess at details that are not clearly visible. Only describe what you can see.
- Use Task 3 preparation time to also plan your Task 4 prediction — both tasks share the same photograph.
CELPIP Speaking – Lesson 9: Task 4 — Making Predictions
Understanding Task 4
Task 4 uses the same photograph as Task 3 and asks you to predict what will probably happen next — what will the people in the image do? What will occur in this scene after the photograph was taken? You have 30 seconds to prepare and 60 seconds to speak. Task 4 tests your ability to use future language accurately, to build logical predictions from visible evidence, and to give reasons for your predictions.
The most important rule for Task 4 is that your predictions must be rooted in visible clues from the photograph — not in random speculation. A prediction without a visual basis does not fulfill the task. A prediction that connects directly to something visible in the photograph and then uses future language to describe what will happen demonstrates both Task Fulfillment and Coherence and Cohesion.
The Three-Part Prediction Frame
Structure your Task 4 response using this three-part frame:
Part 1 — Anchor (10–12 seconds): Reference one specific visible element from the photograph that serves as the basis for your first prediction. "Looking at this scene, I can see [specific detail]. Based on this, I think that [first prediction]."
Part 2 — Two Predictions with Reasons (35–40 seconds): Give two specific, logically grounded predictions with a reason for each. Each prediction should connect to something visible in the photograph. Use hedged future language throughout — not certain future language. Predictions presented as certainties sound unnatural and unrealistic; hedged predictions sound thoughtful and natural.
Part 3 — Broader Outcome (8–10 seconds): End with a brief overall outcome or consequence — what will the situation look like once the predicted events have unfolded. This provides a natural conclusion and demonstrates the ability to project beyond the immediate next moment.
Future Language for Task 4
Task 4 specifically requires you to demonstrate command of future forms. Use a variety of these structures rather than relying entirely on "will."
| Future Form | When to Use It | Example |
|---|---|---|
| will + verb | Confident prediction based on clear evidence | "They will probably start serving the food shortly." |
| is / are going to + verb | Imminent action, clear intention visible in the photo | "The woman looks like she is going to make an announcement." |
| might / may / could + verb | Less certain prediction | "They might decide to move the meeting outside." |
| is likely to + verb | Moderate confidence prediction | "The group is likely to wrap up their discussion soon." |
| I would expect + noun/verb | Thoughtful inference based on context | "I would expect some questions to follow the presentation." |
| Once / After / When + [event], [result] | Sequencing future events | "Once they finish setting up, the event will probably begin." |
Full 60-Second Template for Task 4
Anchor (10–12 seconds): "Looking at this photograph, I can see [specific visible element]. This tells me that [immediate context]. Based on this, I would predict that [first prediction]."
Prediction 1 with reason (18–20 seconds): "Specifically, I think [prediction 1] is likely to happen because [visible reason from the photograph]. [One additional sentence elaborating on this prediction]."
Prediction 2 with reason (18–20 seconds): "I would also expect that [prediction 2]. The reason I say this is [visible or logical basis]. This could lead to [brief consequence]."
Broader outcome (8–10 seconds): "Overall, I think the situation will probably [broader outcome or resolution]. It looks like [final statement about the scene's direction]."
Worked Example: Complete Task 4 Response
Photograph described in Task 3: A group of four people seated around a table in what appears to be an office meeting room. They have papers and laptops in front of them. One person is standing and appears to be presenting something. The atmosphere looks focused.
Task 4 response: "Looking at this photograph, I can see a group of colleagues in what appears to be an active work meeting, with one person standing and presenting. Based on this, I would predict that the presentation is about to wrap up and the group will move into a discussion phase. Specifically, once the presenter finishes their points, the seated team members will probably ask questions or share their own perspectives. I notice that a couple of them have notes in front of them, which suggests they came prepared to contribute — so I would expect a fairly lively back-and-forth conversation to follow. I would also predict that by the end of the meeting, the group will likely assign specific action items or next steps to each person. Meetings at this stage of a project tend to end with a clear division of responsibilities, and the focused atmosphere in the room suggests that decisions are being made. Overall, I think the meeting will probably conclude productively, with everyone leaving with a clear understanding of what they need to do next."
Key Takeaways from Lesson 9
- Task 4 predictions must be based on visible clues from the photograph — not on random ideas.
- Use a variety of future forms: will, going to, might, may, is likely to, would expect.
- Give two predictions, each with a reason connected to something visible in the photograph.
- Use hedged future language — predictions presented as absolute certainties sound unnatural.
- End with a broader outcome that projects the scene's likely resolution.
- Use Task 3 preparation time to identify your Task 4 prediction clues — this saves preparation time in Task 4.