5.2 Assessment Task Types, Design & Self/Peer Evaluation
Key Takeaways
- Objective assessment tasks have a single correct answer and high marker reliability, whereas subjective assessment tasks rely on professional judgment and qualitative rubrics.
- Direct testing tasks require learners to perform the actual skill being assessed, while indirect testing measures underlying knowledge through isolated linguistic sub-skills.
- High-quality assessment balances validity (measuring what it intends to measure), reliability (consistency of results), and practicality (feasibility of administration and marking).
- Self-assessment and peer evaluation promote learner autonomy, metacognitive awareness, and active engagement with learning criteria.
- Effective rubric design employs clear band descriptors, standardized scoring criteria, and moderation procedures to maximize inter-rater and intra-rater reliability.
5.2 Assessment Task Types, Design & Self/Peer Evaluation
Core Concept: Designing effective language assessments requires selecting appropriate task formats that align with learning objectives while balancing validity, reliability, and practical constraints. Furthermore, integrating learner-centered evaluation methods—such as self- and peer-assessment—empowers students to monitor their own linguistic development.
Categorizing Assessment Task Types
Language assessment tasks are classified according to how they are scored (objective vs. subjective) and how directly they measure target communicative skills (direct vs. indirect).
1. Objective vs. Subjective Assessment Tasks
Objective Assessment Tasks
Objective tasks have predetermined, unambiguous correct answers. Because scoring does not require subjective judgment, any marker following the answer key will arrive at the exact same score.
Common Objective Formats:
- Multiple-Choice Questions (MCQs): Selecting one correct option from multiple distractors.
- Matching Tasks: Pairing words with definitions, pictures, or synonyms.
- Gap-Fill / Cloze Exercises: Inserting the correct word or phrase into a text gap.
- True / False / Not Given: Evaluating statements against a reading or listening passage.
Advantages & Disadvantages:
- Pros: Extremely high marker reliability, fast and automated scoring, suitable for large-scale testing.
- Cons: Susceptible to guessing, difficult to write plausible distractors, cannot directly test productive language generation (speaking/writing).
Subjective Assessment Tasks
Subjective tasks require test-takers to produce open-ended responses that must be evaluated by a marker using professional judgment and established scoring rubrics.
Common Subjective Formats:
- Extended Writing Prompts: Essays, letters, report writing, short stories.
- Oral Interviews & Role-Plays: Pair discussions, individual presentations, communicative interactions.
- Portfolio Assessment: Collections of student work compiled over time.
Advantages & Disadvantages:
- Pros: High content and construct validity, measures authentic productive communication and creative language use.
- Cons: Lower marker reliability if rubrics are vague, time-consuming to grade, potential for marker bias (e.g., halo effect).
2. Direct vs. Indirect Testing
| Dimension | Direct Testing | Indirect Testing |
|---|---|---|
| Definition | Assesses language performance by requiring candidates to perform the actual skill being tested. | Assesses underlying language knowledge or sub-skills thought to underpin a skill. |
| Classroom Example | Testing writing by having students write an email; testing speaking via an interview. | Testing writing by asking students to identify grammar errors in pre-written sentences. |
| Validity | High Construct Validity: Directly mirrors real-world language tasks. | Lower Direct Validity: Measures knowledge about language rather than performance. |
| Reliability & Practicality | Requires subjective scoring rubrics; high marking time. | Highly objective; quick and easy to score automatically. |
3. The Assessment Quality Triad: Validity, Reliability & Practicality
When designing or choosing assessment tasks, ELT professionals must navigate the balance between three core principles:
A. Validity
Validity refers to the extent to which an assessment measures what it is intended to measure and nothing else.
- Content Validity: Does the test adequately sample the full scope of the curriculum or skill domain?
- Construct Validity: Does the test accurately reflect underlying theoretical constructs (e.g., communicative competence)?
- Face Validity: Does the test appear fair, relevant, and credible to candidates and stakeholders?
B. Reliability
Reliability concerns the consistency and stability of test results across different testing conditions, times, and markers.
- Test-Retest Reliability: Consistency of results if the same candidate takes the test on different occasions.
- Inter-Rater Reliability: Degree of agreement between two or more independent markers evaluating the same performance.
- Intra-Rater Reliability: Consistency of a single marker's scoring across multiple papers or over time.
C. Practicality
Practicality relates to the realistic constraints of test administration, including financial cost, time required for preparation and marking, facility requirements, and staffing resources.
The Assessment Trade-Off: Increasing directness and validity (e.g., 30-minute individual speaking tests) often reduces practicality and marker reliability. Conversely, highly practical objective tests may sacrifice direct validity for high reliability.
Task Types Comparison Table
| Task Format | Objective / Subjective | Direct / Indirect | Primary Skill / Sub-skill Tested | Marker Reliability | Practicality |
|---|---|---|---|---|---|
| Multiple Choice | Objective | Indirect | Listening / Reading comprehension, Grammar, Vocabulary | Extremely High | High |
| Gap-Fill (Fixed) | Objective | Indirect | Structural accuracy, Collocations, Spelling | High | High |
| Guided Essay | Subjective | Direct | Productive Writing, Text Coherence, Register | Moderate to High (with rubric) | Moderate |
| Paired Role-Play | Subjective | Direct | Interactive Speaking, Fluency, Pragmatics | Moderate (requires training) | Low to Moderate |
4. Self-Assessment and Peer Assessment
Modern ELT practices emphasize involving learners actively in the evaluation process to foster learner autonomy and metacognitive reflection.
Self-Assessment
Self-assessment involves learners evaluating their own language performance, progress, and strategy use against explicit criteria or "Can-Do" statements (e.g., CEFR self-assessment grids).
Pedagogical Benefits:
- Develops self-monitoring and metacognitive strategies during language production.
- Encourages ownership of learning and helps set realistic personal goals.
- Reduces test anxiety by shifting focus to personal growth.
Peer Assessment
Peer assessment involves learners evaluating each other's work using shared criteria, checklists, or scoring rubrics.
Pedagogical Benefits:
- Enhances critical reading and listening skills as students analyze peers' language output.
- Promotes collaborative learning and collaborative feedback loops.
- Demystifies assessment criteria by requiring learners to apply rubrics actively.
Best Practices for Implementing Self & Peer Evaluation:
- Provide Clear Criteria: Use simplified, student-friendly rubrics or checklists rather than vague prompts.
- Train Learners: Model how to give constructive, specific feedback using "Praise, Question, Suggest" protocols.
- Combine with Teacher Feedback: Use self and peer evaluation as formative stepping stones prior to final teacher submission.
Practical Classroom Scenarios
Scenario A: Improving Writing Reliability
A Director of Studies notices that two writing examiners give vastly different scores to the same student essays. To resolve this, the school introduces a detailed analytic rubric with clear band descriptors for Content, Communicative Achievement, Organization, and Language, followed by a moderation workshop.
- Assessment Principle: Enhancing Inter-Rater Reliability. Standardizing rubrics and conducting marker calibration ensures consistent scoring across different raters.
Scenario B: Implementing Peer Feedback in Writing
Before submitting a final draft of a formal letter, students swap drafts and use a checklist to verify whether their partner included an appropriate opening salutation, clear paragraphing, and formal sign-off phrasing.
- Assessment Form: Peer Assessment. Students apply explicit task criteria formatively to assist each other in refining their productive output.
Which assessment task format is classified as an objective task format with high marker reliability?
A language teacher designs a test where students are asked to correct grammatical errors in isolated, pre-written sentences to measure their writing ability. Which descriptor best characterizes this test format?
An examination board conducts a calibration session where multiple markers grade the same set of sample speaking test recordings using a standardized analytic rubric to ensure their scoring is consistent. Which quality of assessment is being strengthened?
What is the primary pedagogical benefit of incorporating learner self-assessment using 'Can-Do' descriptors into a language course?
You've completed this section
Continue exploring other exams