1.2 Computer-Adaptive Multi-Stage Testing (ca-MST) & Rasch Scoring

Key Takeaways

  • The EF SET is calibrated and scored with the Rasch model, in which a single item-difficulty parameter (b) and person ability (theta) sit on the same scale — there is no separate discrimination or guessing parameter.
  • Delivery uses computer-adaptive multi-stage testing (ca-MST) in a 1-3-4 panel: Stage 1 is one medium routing module targeted at mid-B1, Stage 2 offers easy, medium and difficult modules targeted at mid-A2, mid-B1 and mid-B2, and Stage 3 offers four modules targeted at the A1/A2, A2/B1, B1/B2 and B2/C1 cut scores.
  • Routing is by number of questions answered correctly: Stage 1 performance selects the Stage 2 module, and cumulative Stage 1 plus Stage 2 performance selects the Stage 3 module.
  • Twelve routes through the panel are theoretically possible but EF enables only six (M1-M2-M5, M1-M2-M6, M1-M3-M6, M1-M3-M7, M1-M4-M7 and M1-M4-M8), and your reported score comes from all three stages combined.
  • EF caps the standard error of measurement at 0.45 logits on any route of the 50-minute EF SET, and at 0.29 on the 120-minute EF SET PLUS.
Last updated: September 2026

Computer-Adaptive Multi-Stage Testing (ca-MST) & Rasch Scoring

Quick Summary: The EF SET 50 does not calculate your score by totalling correct answers. It is delivered as a computer-adaptive multi-stage test (ca-MST) and calibrated with the Rasch item-response model. Every examinee walks a three-stage route through a pre-assembled panel: one common routing module, then one of three Stage 2 modules, then one of four Stage 3 modules. Your reported score reflects the difficulty of the items you answered correctly, not the percentage you got right.


Psychometric Paradigms: Classical Test Theory vs. Item Response Theory

Traditional classroom and legacy examinations rely on Classical Test Theory (CTT). In a CTT framework:

  • Every test taker encounters an identical linear sequence of questions.
  • Every question is weighted equally (for example, 1 point per correct answer).
  • The final score is a simple raw sum or percentage: $\text{Score} = \frac{\text{Correct Items}}{\text{Total Items}} \times 100$.

The primary flaw of CTT is that score meaning depends entirely on test difficulty. If a form happens to contain unusually difficult reading passages, every candidate's score drops even though their underlying English proficiency has not changed. Conversely, an easy form inflates weak candidates.

EF used CTT during development — its technical report describes classical distractor analysis of every response option — but the operational EF SET is built on Item Response Theory (IRT). Under IRT:

  • Scores locate a candidate on a continuous latent ability scale, written as theta ($\theta$).
  • Each item carries calibrated psychometric properties established by pretesting on hundreds of thousands of learners at all CEFR levels.
  • The probability of a correct response is a mathematical function of both the candidate's ability and the item's characteristics.

The Model EF Actually Uses: Rasch

EF's Academic and Technical Development Report is explicit on this point. Trial items were "subjected to both Classical Test Theory (CTT) and Rasch model analyses," and the Rasch model was adopted "as the psychometric approach to EF SET." EF's stated reason is invariance: "the difficulty of test items must be invariant across test takers … an item intended to measure at the B1 CEFR level must always measure at that level, regardless of the language proficiency level of the test taking population."

The Rasch model has exactly one item parameter — difficulty:

P(Xni=1)=e(θnbi)1+e(θnbi)P(X_{ni} = 1) = \frac{e^{(\theta_n - b_i)}}{1 + e^{(\theta_n - b_i)}}

SymbolNameMeaning inside the EF SET engine
$\theta_n$ (theta)Person abilityCandidate n's latent receptive proficiency. EF reports theta on a scale with a mean of 0 and a practical range of roughly $-3$ to $+3$, then transforms it linearly onto the 0–100 EF scale.
$b_i$Item difficultyThe ability level at which a candidate has a 50% chance of answering item i correctly. Low $b$ items test concrete A1/A2 vocabulary and simple syntax; high $b$ items test C1/C2 idiom, pragmatic nuance and dense rhetorical structure.
Same unitsLogitsRasch places person ability and item difficulty on one common scale, which is what makes it possible to add newly written items to the pool by anchoring them to already-calibrated items.

[!IMPORTANT] Rasch is not the 2PL or 3PL model. Some testing programmes use two- and three-parameter logistic models, which add a discrimination term ($a$) and a pseudo-guessing lower asymptote ($c$). EF SET does not. Under Rasch, every item is assumed to discriminate equally and no guessing floor is modelled, which is precisely why EF can keep expanding its item pool through anchored recalibration without disturbing the reported score scale. If a study resource tells you the EF SET runs a 3PL engine, it is describing a different test.

How Items Enter the Pool

  1. New items are written to a task model for a target CEFR level.
  2. They are administered alongside already-calibrated anchor items.
  3. Rasch analysis places the new items on the same logit scale as the anchors.
  4. Only then can they be assembled into an operational module.

EF recalibrated the entire item pool alongside its July 2014 standard-setting study, which fixed the CEFR cut scores that govern the reported scale to this day.


The 1-3-4 Panel: How EF SET Routes You

Pure item-level adaptive testing recalculates ability after every single response. That is a poor fit for language assessment, because reading a multi-paragraph text or listening to a three-minute lecture only makes sense as a block. EF therefore uses computer-adaptive multi-stage testing, in which items are pre-assembled into modules and modules are pre-assembled into panels.

Each EF SET examinee is administered a three-stage 1-3-4 panel — one module at Stage 1, a choice of three at Stage 2, and a choice of four at Stage 3:

Stage / ModuleTarget difficultyMeasures best atTasks per module (EF SET)Tasks per module (EF SET PLUS)
Stage 1 / Module 1MediumMid-B11 reading – 1 listening2 – 2
Stage 2 / Module 2EasyMid-A21 – 12 – 2
Stage 2 / Module 3MediumMid-B11 – 12 – 2
Stage 2 / Module 4DifficultMid-B21 – 12 – 2
Stage 3 / Module 5Very easyA1/A2 cut1 – 13 – 3
Stage 3 / Module 6Easy–mediumA2/B1 cut1 – 13 – 3
Stage 3 / Module 7Medium–difficultB1/B2 cut1 – 13 – 3
Stage 3 / Module 8Very difficultB2/C1 cut1 – 13 – 3

A "task" here is a whole stimulus with its question set — a reading text with its cluster of items, or a recording with its items — which is why a section built from three modules can still deliver roughly 50 questions.

The Routing Rules

  1. Stage 1 is common to everyone. Module 1 is a medium module targeted at the mid-point of B1. Your ability is estimated initially by the number of questions you answer correctly on Module 1, compared against thresholds embedded in the panel's data-management software.
  2. Stage 1 selects your Stage 2 module — easy (mid-A2), medium (mid-B1) or difficult (mid-B2).
  3. Cumulative Stage 1 + Stage 2 performance selects your Stage 3 module — one of four modules built to measure most precisely at the CEFR boundaries themselves.
  4. All three stages count. EF states that "the total EF SET reading or listening score is the result of performance on all three stages." Stage 1 is not a throwaway warm-up, and Stage 3 is not a tiebreaker.

Only Six of Twelve Routes Are Live

Twelve paths through a 1-3-4 panel are arithmetically possible, but EF enables only six, disabling the paths its data shows to be improbable:

M1 ──> M2 (Easy) ────> M5 (Very easy, A1/A2 cut)
   │            └────> M6 (Easy-medium, A2/B1 cut)
   ├──> M3 (Medium) ─> M6 (Easy-medium, A2/B1 cut)
   │            └────> M7 (Medium-difficult, B1/B2 cut)
   └──> M4 (Difficult) ─> M7 (Medium-difficult, B1/B2 cut)
                └────────> M8 (Very difficult, B2/C1 cut)

The practical reading of this diagram is that the engine will not swing you from the easiest Stage 2 module straight to the hardest Stage 3 module in a single step. A sharp reversal of performance is absorbed gradually, over two routing decisions, rather than in one leap.

Measurement Precision

EF constrains panel assembly so that the standard error of measurement on any enabled route does not exceed 0.45 on the 50-minute EF SET (and 0.29 on the 120-minute EF SET PLUS). That difference is the honest reason the longer test exists: more tasks per module means a tighter confidence interval around your theta estimate.


The Raw-Score Paradox: An Illustration

The following walkthrough uses illustrative figures — EF does not publish per-module item counts — to show why raw percentages cannot be compared across routes.

  • Candidate Alpha performs strongly on Module 1, routes to the difficult Stage 2 module (mid-B2) and then to Module 8 (B2/C1 cut). Facing dense academic prose, Alpha answers 60% of Stage 2 and Stage 3 items correctly.
  • Candidate Beta struggles on Module 1, routes to the easy Stage 2 module (mid-A2) and then to Module 5 (A1/A2 cut). Facing short notices and routine phrases, Beta answers 87% correctly.
DimensionCandidate AlphaCandidate Beta
Route through the panelM1 → M4 → M8M1 → M2 → M5
Stage 2/3 target difficultyMid-B2, then the B2/C1 cutMid-A2, then the A1/A2 cut
Raw accuracy after Stage 160%87%
Where the estimate settles ($\theta$)High positive logitsLow negative logits
Reported EF SET score (0–100)Upper bandLower band
Reported CEFR levelC1 / C2 territoryA1 / A2 territory

Why This Happens

Under Rasch, the evidence a response carries depends on the gap between $\theta$ and $b$:

  • Alpha's correct answers land on items with high difficulty. Each one is improbable for a low-ability candidate, so each contributes strong evidence of high ability.
  • Beta's correct answers land on items with very low difficulty. Getting them right merely confirms Beta is not a complete beginner; getting one wrong is far more informative, because it is statistically surprising.

The lesson is not that you should try to answer harder questions. It is that accuracy on the easy items early in a section carries disproportionate weight, because those items decide which route you walk for the remaining two stages.


What Adaptivity Means for How You Answer

EF does not publish a question-review or back-navigation policy for the test client, so treat the following as the reliable operating assumptions:

  1. Between-stage routing is irreversible. Once a module is submitted, the engine has committed you to a branch of the panel. Nothing you do later reopens that decision. This part is documented: routing is computed from performance on the completed module.
  2. Answer every item as though it were final. Never leave an item blank hoping to return to it. There is no negative marking, so an eliminated-then-guessed answer is strictly better than a blank.
  3. Do not bank on a review screen. Build your pacing plan around resolving each item once, in place. Section 7.1 turns this into concrete per-item time budgets.
  4. Your section clock is the real constraint. Individual tasks are untimed; the 25-minute section timer is what actually ends the section.

[!IMPORTANT] Operational advice: If an item looks intractable, systematically eliminate the options you can rule out, select the strongest remaining choice, and advance. Time spent wrestling with one item is taken directly from the items that decide your Stage 3 routing.

Loading diagram...
EF SET ca-MST 1-3-4 Panel and Enabled Routes
Test Your Knowledge

Which item-response model does EF state that it uses to calibrate and score the EF SET?

A
B
C
D
Test Your Knowledge

How is the EF SET panel structured, and how is a test taker routed through it?

A
B
C
D
Test Your Knowledge

Candidate Alpha is routed to the difficult Stage 2 module and answers 60% of the remaining items correctly. Candidate Beta is routed to the easy Stage 2 module and answers 87% correctly. Why can Alpha still finish with a much higher EF SET score and CEFR level?

A
B
C
D
Test Your Knowledge

Which statement about navigation and routing on the EF SET is supported by what EF publishes?

A
B
C
D