18.4 Standard Setting and the Angoff Calculation
Key Takeaways
Angoff panel ratings estimate a minimally competent candidate's item success probability.
Thresholds need a documented candidate definition, evidence and decision rule.
The five-item example estimates 3.96 correct; an at-least rule requires four whole correct items.
Standard-Setting Methodologies & The Modified Angoff Walkthrough
A passing threshold should be supported by the purpose and performance requirements, rather than selected only because a familiar percentage is convenient. Employment-testing requirements depend on the use and context; they do not mandate the Angoff method for every workplace course.
The Modified Angoff method is one criterion-referenced standard-setting approach. A panel estimates the probability that a minimally competent candidate would answer each knowledge item correctly. Panel judgments require a clear competency definition, evidence and review; the method does not guarantee legal defensibility.
The Minimally Competent Candidate (MCC) Concept
The Angoff process hinges upon the concept of the Minimally Competent Candidate (MCC) (also known as the "borderline candidate"). The MCC represents a hypothetical trainee who possesses just enough knowledge, judgment, and skill to perform the job safely and effectively without endangering themselves or others. They are neither a master technician nor a reckless novice; they represent the threshold of acceptable professional competence.
Standard-setting workflow
Define the minimally competent candidate, recruit an appropriate panel, train members on the task, obtain independent ratings, discuss discrepancies and review evidence. Aggregate ratings according to the chosen method and document the final decision. Panel size and rounds depend on the assessment; a fixed six-to-twelve-person rule is not universal.
Angoff Standard-Setting Calculation Walkthrough
| Knowledge item | SME 1 | SME 2 | SME 3 | SME 4 | SME 5 | Mean |
|---|---|---|---|---|---|---|
| 1 | 0.85 | 0.90 | 0.80 | 0.85 | 0.90 | 0.86 |
| 2 | 0.95 | 0.95 | 0.90 | 0.90 | 0.95 | 0.93 |
| 3 | 0.70 | 0.65 | 0.75 | 0.70 | 0.70 | 0.70 |
| 4 | 0.65 | 0.70 | 0.60 | 0.65 | 0.65 | 0.65 |
| 5 | 0.80 | 0.85 | 0.80 | 0.80 | 0.85 | 0.82 |
The five item means sum to 3.96, or 79.2% of five items. If a policy requires at least this many correct whole items, the threshold is four correct answers out of five (80%), because 3.96 items cannot be earned on a dichotomously scored test. This is a knowledge-test illustration, not a method for averaging away critical performance failures.
Retesting Protocols and Record Integrity
Provide targeted remediation and an appropriate reassessment under the applicable program. Define retest conditions, evaluator responsibilities and restrictions before administration. Document the evidence and decision. Do not invent a universal prohibition on same-day retesting or treat a written retest as proof of a missing physical skill.
Reviewing a standard-setting decision
Panel ratings are judgments about a defined minimally competent candidate, not the observed percentage correct from one class. Define that candidate's expected training, duties and performance carefully. Panelists who imagine very different audiences can produce disagreement that averaging will conceal. Discussion should identify those different assumptions before the final decision.
Consider what the threshold permits. A five-item knowledge test offers only six possible raw scores: zero through five. Converting an estimate of 3.96 into a whole-item rule requires an explicit policy; four correct is the first whole score meeting at least 3.96. A percentage displayed with decimal precision does not create fractional credit where every item is dichotomous.
| Standard-setting element | Documentation |
|---|---|
| Candidate definition | Expected minimum competence and conditions |
| Item judgments | Independent estimates and review process |
| Evidence | Content review and relevant performance data |
| Decision | Final threshold, rounding rule and rationale |
Standard setting cannot rescue an invalid assessment. If the items do not represent the objectives, a carefully calculated threshold still measures the wrong content. If the key is incorrect, the score is distorted. If a physical skill is assessed only by written recall, the instrument is insufficient for that outcome. Validate content and administration before interpreting the cut score.
Different purposes may call for different standard-setting approaches. A required procedural criterion may be directly specified by a controlling rule or approved task standard. A panel-estimation method is a useful concept for knowledge testing, but must not average away a safety-critical action. Keep these decision rules distinct.
After use, review the assessment's evidence and decisions. An unexpected failure pattern can reflect a true learning gap, a poor criterion, an item defect or a new audience. Investigate and document changes. Do not lower a threshold automatically because a class dislikes its results, or raise it to create a desired failure quota. The threshold should remain connected to the required competence.
Avoiding false precision
A panel mean such as 79.2% can look more precise than the judgment process supports. Record assumptions, disagreement and the final decision rule. Do not imply that the decimal alone proves a scientifically exact competence boundary.
Standard-setting documentation helps a later reviewer understand how the threshold was chosen and whether changed objectives require reconsideration. It does not exempt the team from validating the items, administration and scoring. The entire evidence chain matters for a sound decision.
Separating threshold review from learner remediation
If many learners narrowly miss a threshold, first examine the evidence. Was the task taught and practiced appropriately? Were the items accurate and representative? Were conditions comparable and accessible? Was scoring correct? An unexpected result warrants investigation, not an automatic threshold change. A confirmed instrument defect needs fair correction; a genuine learning gap needs remediation.
Keep the decision rationale distinct from the remediation record. The standard-setting file explains what level of evidence is required. The learner record explains which criteria were met, which were not and what happened afterward. Mixing these can make it appear that a threshold was changed to pass a particular person.
For performance assessment, the approved critical-step rule remains separate from the knowledge-test calculation. A person scoring four of five on the illustrative written test may meet its raw-score threshold while still lacking a required physical check. Conversely, a successful observed task does not necessarily answer every knowledge requirement. Use the evidence required by the program rather than substituting one favorable measure for another.
Review the standard when the role, objective or criterion changes. Preserve earlier versions and the reason for revision under the record policy. This lets future reviewers understand which standard applied to a particular decision and avoids applying a new threshold retroactively without an appropriate process.
Key takeaways
- Angoff panel ratings estimate a minimally competent candidate's item success probability.
- Thresholds need a documented candidate definition, evidence and decision rule.
- The five-item example estimates 3.96 correct; an at-least rule requires four whole correct items.
Five dichotomous items have an estimated threshold of 3.96 correct. Under an at-least rule, what raw score first meets it?
Three correct items
A fractional score of 3.96 without partial-credit scoring
Four correct items
All five items because every knowledge threshold must be 100%
Sections you finish are checked off in the contents.