4.5 Assessment Methods: Self, Peer, Portfolio, and Continuous Assessment

Key Takeaways

  • The assessment agent is who judges — teacher, self, peer or external body — and is an independent variable from assessment purpose.
  • Self-assessment uses can-do statements, learner diaries and reflection sheets to build autonomy, but is too inaccurate to decide a grade on its own.
  • Peer assessment only works when learners are trained in the procedure and given clear criteria in learner-friendly language.
  • A portfolio assesses work gathered over time — drafts, recordings, projects — as a body of evidence, unlike a single one-off test.
  • A usable instrument needs validity, reliability, practicality and positive washback; objective items score higher on reliability, subjective tasks need rating scales and standardisation.
Last updated: July 2026

4.5 Assessment Methods: Self, Peer, Portfolio, and Continuous Assessment

TKT Module 1 Part 3 asks not only why an assessment happens but who carries it out and how the evidence is collected. Section 4.4 covered assessment purpose; this section covers the assessment agent, the method, and the four qualities that decide whether an instrument is worth using.

Agent vs. Purpose

The agent is whoever makes the judgement: the teacher, the learner (self), a classmate (peer), or an external body such as Cambridge English. Agent and purpose are independent variables — a learner can self-assess formatively with a can-do checklist, while an external body assesses summatively through a B1 Preliminary certificate. TKT items usually describe an activity and ask you to label it: identify the agent first, then the method.

Self-assessment

In self-assessment the learner judges their own performance against stated criteria. The standard instruments are can-do statements (CEFR-style descriptors such as "I can order a meal in a restaurant", rated yes / with help / not yet), learner diaries, and end-of-unit reflection sheets. Example: after a hotel-booking role-play, A2 learners tick a five-point checklist — "I asked about the price", "I used Could I…? twice" — and write one goal for next week.

Self-assessment builds learner autonomy, makes lesson aims transparent, trains metacognition and costs almost no class time. Its weakness is accuracy: learners routinely over- or under-rate themselves, beginners cannot judge their own grammar, and cultural expectations push some towards modesty. It feeds formative decisions and should not alone determine a report grade.

Peer assessment

In peer assessment learners assess each other against shared, explicit criteria. Two classic applications: peer editing of a second draft in process writing, using a correction code or a three-point checklist ("Is there a topic sentence? Are the past forms correct?"); and peer feedback checklists in speaking, where a third learner in a group of three holds an observation grid and ticks turn-taking phrase used / eye contact / accurate past simple while the other two do the task.

Peer assessment raises student talking time, supplies alternative models, and sharpens the critical eye learners later turn on their own work. It collapses without two conditions: training (model the procedure once on the board with a sample text) and clear criteria in learner-friendly language. Untrained peers correct spelling only, or write "very good" to protect a friendship. Frameworks such as two stars and a wish keep feedback balanced.

Portfolio assessment

A portfolio is a collection of a learner's work gathered over time and assessed as a single body of evidence: successive drafts of a writing task, audio recordings from month one and month five, a project poster, a reading log, plus the learner's note on why each piece was included. Contrast a one-off test, where a single 60-minute performance on one morning stands for a whole course. Portfolios show process and progress, and suit learners who freeze under time pressure. The European Language Portfolio (ELP), from the Council of Europe, is the standard ELT example: a Language Passport, a Language Biography with can-do self-assessment grids, and a Dossier of work samples. The cost is marking load and poor comparability between collections.

Continuous assessment

Continuous assessment builds the final grade from work produced across the whole course — homework, short quizzes, presentations, projects — instead of one terminal paper. (Section 4.4 covers summative end-of-course testing itself; the contrast here is that continuous assessment spreads the evidence over many occasions.) It lowers exam anxiety and surfaces problems early, but demands disciplined record-keeping and can tempt learners into minimum effort.

Informal Observation vs. Formal Testing

Informal assessment is what a teacher does while monitoring: circulating during a freer-practice information-gap task with an observation grid or tick-list of target exponents, and keeping anecdotal records ("Marta still produces he don't"). Learners often do not realise they are being assessed, so the language sample is natural. Formal testing applies standardised administration and marking conditions — announced date, identical paper, fixed time limit, silence, an answer key or agreed rating scale — which makes results comparable across learners and classes.

MethodAgentTypical instrumentMain strengthMain limitation
Self-assessmentLearnerCan-do checklist, diaryBuilds autonomy, cheapInaccurate as a grade
Peer assessmentClassmatePeer-edit checklist, gridMore models, involvementNeeds training, criteria
PortfolioTeacher (+ learner)Dossier of drafts, recordingsShows progress over timeHeavy to mark
Continuous assessmentTeacherCoursework record, quizzesLow anxiety, early diagnosisRecord-keeping burden
Informal observationTeacherMonitoring notes, anecdotesNatural language sampleUnsystematic
Formal testTeacher / external bodyStandardised paper, rating scaleComparable resultsTime-consuming, stressful

Qualities of a Good Assessment Instrument

Validity means the instrument tests what it claims to test. Negative worked example: a "reading comprehension" test requiring answers in extended written paragraphs is partly invalid, because a learner who understood the passage perfectly but writes weakly scores low — the task measures writing as well as reading. Multiple-choice, matching or short-answer items restore validity. Likewise a listening test using a recording full of unknown low-frequency vocabulary tests lexis, not listening.

Reliability means consistency: the same performance earns the same mark from different markers (inter-rater reliability) and on different occasions. Objective items — multiple choice, true/false, matching — score high on reliability because an answer key removes marker judgement. Subjective tasks such as essays and interviews become reliable only with a published rating scale, marker standardisation and ideally double marking.

Practicality asks whether the institution can actually run the assessment: writing and marking time, cost, equipment, rooms, invigilators, class size. A 15-minute one-to-one oral interview is highly valid for speaking and entirely impractical for 45 learners in a 50-minute period.

Washback (or backwash) is the effect a test has on what and how teachers teach and learners study. Positive washback: an exam with a paired speaking task pushes teachers to run regular pair discussions. Negative washback: a school exam testing only grammar through multiple choice leads teachers to drill gap-fills and abandon speaking, and learners to memorise item types instead of using the language.

Choosing a Method

Match the method to the aim, the learner group and the stage of the course. Diagnostic aims at the start suit informal observation and self-assessment; certification aims at the end require formal, reliable testing. Young learners respond well to portfolios, can-do stickers and observation, and badly to long formal papers; adults on an exam course expect timed mock tests for the positive washback. Mid-course, low-stakes peer and continuous assessment keep the feedback loop running without exam pressure.

Test Your Knowledge

At the end of a unit, a teacher gives A2 learners a sheet of 'I can...' statements such as 'I can order a meal in a restaurant' and asks them to tick yes, with help, or not yet. Which assessment method is this?

A
B
C
D
Test Your Knowledge

A 'reading comprehension' test requires learners to answer each question in an extended written paragraph. Several learners who clearly understood the text score badly because their writing is weak. Which quality of a good assessment instrument is this test failing?

A
B
C
D
Test Your Knowledge

Which of the following is an example of negative washback?

A
B
C
D