10.5 Assessment Instrument Development, Item Analysis, and Validation
Key Takeaways
Instrument development begins with a defined construct, intended use, population, and score interpretation—not with writing attractive items.
A table of specifications links content domains and cognitive demand to the intended score claim.
Item difficulty, discrimination, option functioning, and qualitative review supply different evidence and must be interpreted together.
Reliability and validity concern score use in a population and context; they are not permanent labels attached to a test.
Fairness, accessibility, security, documentation, and revision operate throughout the development cycle.
Assessment Instrument Development, Item Analysis, and Validation
Test construction is a chain of evidence. A polished questionnaire is not automatically a valid instrument, and a high reliability coefficient cannot rescue a score that does not support its intended interpretation. Development starts by stating what decision the score will inform, for whom, and under what conditions.
Define the Construct and Intended Use
A construct is the attribute to be measured, such as career decision self-efficacy, reading comprehension, or counseling knowledge. Write a construct map that describes important facets and, where relevant, increasing levels of performance. Define the target population and exclusions. A scale developed for employed adults cannot be assumed suitable for early adolescents merely because the language seems simple.
Specify whether the result will support screening, diagnosis by an authorized professional, placement, program evaluation, selection, or formative feedback. The higher the stakes, the stronger the evidence and procedural safeguards required. Also state what the score does not mean; this prevents interpretive overreach.
Blueprint Before Writing Items
A table of specifications cross-tabulates content or construct dimensions with cognitive processes or task types. It protects representativeness. If 40% of the intended domain concerns application but only recall items are written, the test underrepresents the construct even if every fact is correct.
Item writers then choose formats suited to the claim: selected response, constructed response, rating scale, performance task, interview, or observation. Multiple-choice items need one defensible best answer, plausible distractors, relevant information, and freedom from unintended clues. Avoid grammatical mismatches, unequal option length that signals the key, unnecessary negatives, and cultural knowledge unrelated to the construct. Rating-scale items should express one idea at a time and provide response categories that are meaningful and ordered.
Expert Review and Cognitive Pretesting
Subject-matter experts review content relevance, accuracy, representativeness, bias, and intended cognitive level. Accessibility reviewers examine language, layout, disability access, and avoidable construct-irrelevant barriers. Cognitive interviews or think-aloud trials explore how members of the target population understand instructions and items. If respondents choose the keyed answer for a reason unrelated to the intended construct, statistical performance alone may conceal a defect.
Pilot Testing and Classical Item Analysis
Pilot the draft under conditions resembling intended administration and record changes. For a dichotomously scored item:
- Difficulty index (p) is the proportion answering correctly. A p of .80 means the item was easier for that sample than an item with p of .35; it does not mean 80% “validity.”
- Discrimination indicates how well performance on the item distinguishes examinees with stronger versus weaker performance on the relevant overall measure. A negative value is a warning: possible miskey, ambiguity, content mismatch, or unusual subgroup behavior.
- Distractor analysis checks whether incorrect options attract some examinees, particularly those with weaker mastery. A never-chosen distractor contributes little.
For norm-referenced selections, items near moderate difficulty often provide more differentiation, but the desired difficulty depends on purpose. A mastery test may appropriately contain many items that competent learners answer correctly.
Reliability and Measurement Precision
Estimate reliability using a method that matches the design: internal consistency for item homogeneity, test-retest for temporal stability, alternate forms for form equivalence, and inter-rater evidence for scored performances. Report the standard error of measurement so users understand score precision. Very high internal consistency can sometimes indicate redundant items rather than broad coverage.
Reliability is necessary for many uses but does not prove validity. A bathroom scale that consistently adds five kilograms is reliable but inaccurate; similarly, a test can consistently measure the wrong or too-narrow attribute.
Build a Validity Argument
Modern validity practice accumulates evidence supporting the proposed interpretation and use:
| Evidence source | Core question |
|---|---|
| Test content | Do tasks adequately represent the domain? |
| Response processes | Are respondents and raters using the intended processes? |
| Internal structure | Do item relationships fit the proposed dimensions? |
| Relations to other variables | Are expected convergent, discriminant, or criterion patterns present? |
| Consequences and fairness | What intended and unintended effects follow from use? |
Evidence is use-specific. Correlation with college grades may support one predictive use but not clinical diagnosis. Translation requires more than word substitution; adaptation should examine construct equivalence, language, differential performance, and new evidence in the Philippine population.
Norming, Standard Setting, and Documentation
If norm-referenced interpretation is intended, the normative sample must represent the relevant population and the sampling date must be reported. Criterion-referenced decisions may require a defensible standard-setting method rather than a norm. Document administration, scoring, accommodations, evidence, limitations, security, and revision history in a technical manual.
Revise and Monitor
Remove or repair defective items only after considering content coverage; dropping every statistically weak item can narrow the construct. After operational use, monitor exposure, drift, subgroup fairness, changes in the population, and consequences. Revalidation is required when use, language, mode, or population changes materially.
Exam Application
On an item-analysis vignette, diagnose before deleting. Check the key, wording, blueprint match, option functioning, sample, and subgroup patterns. On a validity question, identify the proposed score use and ask which evidence actually supports that claim.
A multiple-choice item has a negative discrimination index. What is the best initial interpretation?
The item is automatically the best item on the test
The item may be miskeyed, ambiguous, mismatched to the domain, or behaving unusually and should be investigated
The test has perfect validity
The item must be retained without review because it is difficult
Why is a table of specifications prepared before final item writing?
To guarantee a high pass rate
To replace all expert review
To link domain content and cognitive demand to the intended score interpretation
To keep every item at identical difficulty
A career scale has high internal consistency in a U.S. adult sample. What can a Philippine school counselor conclude about using it with Grade 8 learners?
It is automatically valid for diagnosis in every population
Translation alone makes the evidence transferable
Reliability proves that the construct is culturally identical
The reported coefficient is insufficient; the intended use, adaptation, population, fairness, and local validity evidence must be examined
Sections you finish are checked off in the contents.