9.2 Validity Types, Test Standardization, and Norm vs. Criterion Referencing

Key Takeaways

  • Validity is not an intrinsic property of a test instrument, but the degree to which empirical evidence and theoretical rationales justify the specific interpretations and uses of test scores.

  • While classical psychometrics conceptualized validity through the tripartite division of content, criterion, and construct validity, contemporary professional testing standards treat validity as a unitary construct centered on construct validation.

  • Content validity requires systematic domain sampling through a Table of Specifications (TOS); criterion validity examines concurrent or predictive real-world outcomes; and construct validity verifies theoretical trait boundaries through convergent, discriminant, and factor-analytic evidence.

  • Standardization establishes test equity and objectivity by enforcing identical administration conditions, standardized scripts, precise timing, and representative normative sampling.

  • Norm-Referenced Testing (NRT) evaluates an individual's relative rank within a comparative peer group across a bell curve, whereas Criterion-Referenced Testing (CRT) measures mastery against absolute competency cut scores, such as the 75% passing standard on the GCLE.

Last updated: September 2026

9.2 Validity Types, Test Standardization, and Norm vs. Criterion Referencing

Notice: This chapter provides independent study preparation for candidates studying for the Philippine Guidance Counselor Licensure Examination (GCLE). This material is independently developed to support mastery of psychometric validity, test standardization, and appraisal systems.

A psychological or educational test may exhibit flawless reliability—producing consistent, replicable scores with minimal measurement error—yet remain completely useless or misleading if it fails to measure the intended psychological construct. In the hierarchy of psychometrics, validity is the supreme evaluative standard. A guidance counselor who selects, administers, or interprets tests without empirical validity evidence risks misdiagnosing learning challenges, misdirecting career paths, and committing ethical violations under Republic Act No. 9258 and the PRBGC Code of Ethics.


The Concept of Psychometric Validity and Professional Standards

According to the Standards for Educational and Psychological Testing, jointly developed by the American Educational Research Association (AERA), the American Psychological Association (APA), and the National Council on Measurement in Education (NCME):

Core Psychometric Rule: Validity is the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of interpretations and actions based on test scores.

Critical Tenets of Modern Validity Theory

  1. Validity Inheres in Inferences, Not Instruments: Psychometricians do not declare that "this test is valid." Rather, validity pertains to the specific interpretation of a test score for a specific purpose with a defined population. For instance, an English reading comprehension test may possess high validity for predicting academic success in high school humanities, but zero validity for diagnosing clinical depression or selecting candidates for mechanical drafting.
  2. Reliability is Necessary, But Not Sufficient: An instrument cannot be valid unless it is first reliable (unreliable scores represent random error, which cannot correlate with anything meaningful). However, high reliability does not guarantee validity. A miscalibrated tape measure will reliably yield identical lengths every time, but it remains completely invalid for measuring weight or true distance.
  3. Validity is a Matter of Degree: Validity is never an all-or-none dichotomy. A test possesses low, moderate, or high validity evidence for particular operational uses.

The Tripartite Model and Contemporary Unitary Validity

Historically, psychometrics categorized validity into three independent pillars, often referred to as the Tripartite View or "Holy Trinity": Content Validity, Criterion-Related Validity, and Construct Validity (formalized by Cronbach and Meehl in 1955).

In contemporary measurement theory (spearheaded by Samuel Messick and codified in the AERA/APA/NCME Standards), validity is conceptualized as a unitary construct. All validity evidence ultimately serves to support the validity of the construct interpretation. Rather than thinking of separate "types" of validity, modern counselors evaluate multiple lines of evidentiary support.

                         [Unitary Validity Concept]
                                    │
         ┌──────────────────────────┼──────────────────────────┐
         ▼                          ▼                          ▼
[Content-Related Evidence]  [Criterion-Related Evidence]  [Construct-Related Evidence]
- Domain definition         - Concurrent validity          - Convergent validity
- Table of Specifications   - Predictive validity          - Discriminant validity
- Lawshe's CVR              - Incremental validity         - Multitrait-Multimethod
- Subject Matter Experts    - Avoid criterion bias         - Factor Analysis (EFA/CFA)

1. Content-Related Validity Evidence

Content-related validity evaluates how comprehensively and representatively the behaviors, questions, or tasks on a test sample the targeted domain of knowledge, skill, or psychological functioning.

The Table of Specifications (TOS)

Content validity is not determined after a test is printed; it is engineered from inception using a Table of Specifications (TOS). A TOS is a two-way grid that crosses content topics with cognitive process levels (such as Bloom's Revised Taxonomy: Remembering, Understanding, Applying, Analyzing, Evaluating, Creating):

  • It dictates the exact percentage and distribution of test items allocated to each content topic.
  • It prevents construct underrepresentation (failing to include important dimensions of the domain) and construct-irrelevant variance (introducing extraneous factors, like overly complex vocabulary on a math computation exam).

Quantitative Content Validity: Lawshe's CVR

In 1975, C.H. Lawshe established a quantitative method for evaluating content validity using a panel of Subject Matter Experts (SMEs). Each panelist evaluates every test item by answering: "Is the skill or knowledge measured by this item 'essential', 'useful but not essential', or 'not necessary' to the domain?"

The Content Validity Ratio (CVR) for each item is calculated as:

CVR = [n_e - (N / 2)] / (N / 2)

Where:

  • nen_e = number of expert panelists rating the item as "essential"
  • NN = total number of expert panelists on the review board

The resulting CVR ranges from -1.00 (when no panelists deem the item essential) through 0.00 (when exactly half deem it essential) to +1.00 (when all panelists agree it is essential). Items are retained only if their CVR exceeds statistical significance based on Lawshe's minimum cutoff values for a given panel size.

Face Validity vs. True Content Validity

  • Face Validity: The superficial appearance of a test to examinees, parents, or laypersons—whether the test looks like it measures what it claims to measure. Face validity is not a statistical or psychometric property. A test can possess high face validity while being completely invalid, or low face validity (e.g., subtle MMPI personality items) while exhibiting formidable scientific validity. However, face validity is practically vital in school and counseling settings to foster examinee rapport, cooperation, and motivation.
  • True Content Validity: Established through rigorous structural domain analysis, expert panels, and statistical TOS verification.

2. Criterion-Related Validity Evidence

Criterion-related validity evaluates how effectively test scores predict or correlate with an external, independent, and operationalized benchmark or outcome (the criterion).

Concurrent Validity vs. Predictive Validity

The critical psychometric distinction between concurrent and predictive validity is the time dimension (temporal latency) between administering the test and obtaining the criterion measure:

DimensionConcurrent ValidityPredictive Validity
Time FrameTest and criterion collected at the same timeTest administered first; criterion measured in future
Core QuestionDoes the test reflect current diagnostic or behavioral status?Does the test forecast future performance or behavior?
Typical ScenarioValidating a 15-minute depression screener against the 45-minute Beck Depression Inventory administered on the same dayUsing a senior high school College Entrance Test (CET) to predict first-year college Grade Point Average (GPA)
Primary UtilityEfficient diagnostic screening, replacing expensive or invasive batteriesSelection, admissions, academic placement, vocational hiring

Critical Threats to Criterion Validity

  • Criterion Contamination: A severe methodological flaw that occurs when the individual responsible for assigning the criterion rating has prior knowledge of the examinee's predictor test scores. For example, if a high school science teacher knows which students scored in the 99th percentile on a science aptitude test, the teacher's subjective semester grades (the criterion) will be biased by halo effects. The resulting correlation will be artificially inflated, producing spurious validity evidence. Criterion raters must remain blind to predictor scores.
  • Standard Error of Estimate (SEestSE_{est}): While the correlation coefficient (rxyr_{xy}) quantifies the strength of association between test and criterion, the SEestSE_{est} reflects the margin of error in predicting an individual's criterion score from their test score:
SE_est = s_y * sqrt(1 - r_xy^2)

Where sys_y is the standard deviation of the criterion measure and rxyr_{xy} is the criterion validity coefficient.

  • Incremental Validity: The degree to which a newly introduced assessment instrument adds unique predictive power above and beyond existing assessment tools or demographic predictors. If an expensive personality battery increases predictive power of academic performance by only 0.01 beyond standard GPA and aptitude tests, it lacks incremental validity.

3. Construct-Related Validity Evidence

Construct-related validity evaluates the degree to which an instrument measures a complex, theoretical, non-observable psychological construct (e.g., emotional intelligence, self-efficacy, scholastic aptitude, career maturity). Because constructs cannot be touched or directly observed, construct validation requires accumulating evidence from multiple theoretical and empirical directions.

Convergent vs. Discriminant (Divergent) Validity

  • Convergent Validity: Demonstrated when an instrument exhibits substantial positive correlations with other tests, measures, or behaviors that theoretically measure the same or related construct. For example, a newly developed adolescent anxiety inventory should correlate highly (r>0.60r > 0.60) with established anxiety scales like the Revised Children's Manifest Anxiety Scale (RCMAS).
  • Discriminant (Divergent) Validity: Demonstrated when an instrument exhibits low, near-zero, or non-significant correlations with measures of constructs that are theoretically unrelated. For example, the same adolescent anxiety scale should correlate near zero (r≈0.00r \approx 0.00 to 0.150.15) with measures of mechanical spatial reasoning or mathematical calculation. If it correlates strongly with math computation, the anxiety scale is contaminated with construct-irrelevant cognitive variance.

The Multitrait-Multimethod Matrix (MTMM)

Introduced in 1959 by Donald T. Campbell and Donald W. Fiske, the Multitrait-Multimethod Matrix (MTMM) is a sophisticated experimental paradigm for evaluating convergent and discriminant validity simultaneously. The MTMM requires evaluating at least two different traits (e.g., Trait A: Anxiety, Trait B: Aggression) measured by at least two different methods (e.g., Method 1: Self-Report Inventory, Method 2: Peer Behavioral Rating).

                      Trait A1 (Self)  Trait B1 (Self)  Trait A2 (Peer)  Trait B2 (Peer)
Trait A1 (Self-Anx)   [Reliability] 
Trait B1 (Self-Aggr)     r_hetero1      [Reliability]
Trait A2 (Peer-Anx)      CONVERGENT        r_method      [Reliability]
Trait B2 (Peer-Aggr)     r_null          CONVERGENT         r_hetero2      [Reliability]

Four distinct correlation coefficients are analyzed within the matrix:

  1. Monotrait-Monomethod (Reliability Diagonal): Same trait, same method (e.g., Self-Report Anxiety with Self-Report Anxiety). Represents test reliability; must be the highest coefficients in the matrix.
  2. Monotrait-Heteromethod (Validity Values): Same trait, different methods (e.g., Self-Report Anxiety with Peer-Rated Anxiety). Demonstrates convergent validity; must be statistically significant and substantially high.
  3. Heterotrait-Monomethod: Different traits, same method (e.g., Self-Report Anxiety with Self-Report Aggression). Reflects common method bias; should be substantially lower than convergent validity.
  4. Heterotrait-Heteromethod: Different traits, different methods (e.g., Self-Report Anxiety with Peer-Rated Aggression). Demonstrates discriminant validity; must be the lowest values in the matrix.

Factor Analysis

Factor analysis is a multivariate statistical technique used to establish construct validity by examining the underlying latent structure among a large battery of test items or subtests:

  • Exploratory Factor Analysis (EFA): Used during initial instrument development to discover how many latent dimensions (factors) account for item correlations without preconceived theoretical constraints.
  • Confirmatory Factor Analysis (CFA): Used to test whether empirical test data conforms to a specific a priori theoretical model (e.g., confirming whether a career inventory genuinely breaks into John Holland's six RIASEC dimensions).

The Test Standardization Process

Test standardization is the process of establishing uniform procedures for administering, scoring, and interpreting a psychological instrument to ensure equity and comparability across all test-takers.

Essential Requirements of Standardized Testing

  1. Standardized Administration Protocols: The test manual must provide verbatim oral instructions, exact time limits, specified test materials, and defined environmental conditions (lighting, seating, acoustics). Any deviation by a counselor (such as giving extra time or explaining difficult words) invalidates the standardized administration.
  2. Objective Scoring Rubrics: Explicit scoring keys and rubrics ensure that identical responses receive identical scores regardless of who grades the test.
  3. The Normative Sample (Standardization Sample): The test must be administered to a large, representative sample of the target population. Sampling must use stratified random selection matching demographic variables such as age, gender, geographic region, socioeconomic status, and ethnicity.

Norm-Referenced Testing (NRT) vs. Criterion-Referenced Testing (CRT)

One of the most essential distinctions tested on the GCLE is the structural contrast between Norm-Referenced Testing (NRT) and Criterion-Referenced Testing (CRT).

Assessment FeatureNorm-Referenced Testing (NRT)Criterion-Referenced Testing (CRT)
Primary ObjectiveCompare an individual's score against a peer groupDetermine mastery of specific skills or knowledge
Core Question Answered"How does this student's score rank relative to others?""What specific competencies can this student perform?"
Score InterpretationRelative standing (percentiles, z-scores, stanines, T-scores)Absolute standard (percentage correct, pass/fail cut score)
Distribution of ScoresNormal distribution (bell curve) with wide dispersionSkewed distribution (ideally J-shaped as students master skills)
Item Selection StrategySelects items with moderate difficulty (p≈0.50p \approx 0.50) to maximize varianceSelects items that directly measure domain competencies regardless of difficulty
Benchmark / StandardEstablished by performance of the normative sampleEstablished by predefined curriculum objectives or professional standards
Typical ExamplesWechsler Intelligence Scales (WAIS/WISC), 16PF, College Entrance TestsPhilippine Guidance Counselor Licensure Exam (GCLE), driver licensing tests

Criterion Cut-Scores and Standard Setting

In Criterion-Referenced Testing, passing or failing hinges upon a cut-score. Professional licensing boards, including the Professional Regulatory Board of Guidance and Counseling (PRBGC), establish cut-scores using formal psychometric standard-setting methods:

  • The Angoff Method: A panel of Subject Matter Experts reviews each item and estimates the proportion of "minimally competent candidates" who would answer it correctly. The sum of these probabilities forms the passing standard.
  • The Statutory Standard in RA 9258: For the GCLE, Section 17 of Republic Act No. 9258 establishes a strict criterion standard: To pass the examination, a candidate must obtain a General Weighted Average (GWA) of at least seventy-five percent (75%), with no grade below sixty percent (60%) in any single statutory subject area. This standard is purely criterion-referenced; every candidate who attains this benchmark passes, regardless of whether 10% or 90% of the cohort achieves it.

High-Yield Exam Watch: Traps and Clinical Vignettes

Exam Trap Alert: Do not confuse Content Validity with Face Validity:

  • Face Validity: Subjective appearance to laypersons ("Does it look like an anxiety test?"). Affects motivation but has zero statistical standing.
  • Content Validity: Systematic domain representation verified through a Table of Specifications (TOS), expert SME reviews, and Lawshe's CVR.

Clinical Vignette: The Misguided Curve Dilemma

A newly appointed university guidance director discovers that a faculty member graded a final competency-based counseling practicum examination by 'grading on a curve' (forcing student grades into a normal bell-curve distribution with 10% A's, 20% B's, 40% C's, 20% D's, and 10% F's). Several students who demonstrated 90% mastery of core microskills received failing grades simply because they scored in the lower tail of this exceptionally high-performing class. What psychometric violation occurred?

Psychometric Analysis: The faculty member misapplied Norm-Referenced assumptions to a Criterion-Referenced learning domain. Clinical counseling competencies and professional ethics are criterion domains where absolute mastery is required for safe public practice. Forcing scores into a bell curve arbitrarily manufactures artificial failure, penalizing students who demonstrated verified mastery of essential microskills simply because their peers also excelled. The guidance counselor should explain that professional competency assessment mandates criterion-referenced cut-scores, where grading reflects absolute attainment against predefined rubrics rather than relative peer ranking.

Loading diagram...
Contemporary Unitary Validity Framework and Assessment Typologies
Test Your Knowledge

An industrial guidance counselor uses a spatial reasoning test to screen applicants for a drafting program. At the end of the semester, the drafting instructor rates the job performance of each student. However, the instructor was provided with the students' entrance reasoning test scores before assigning final performance ratings. The correlation between test scores and instructor ratings was unusually high at r = 0.82. What psychometric artifact threatens the validity of this criterion evaluation?

A

Construct underrepresentation

B

Range restriction in the predictor

C

Regression to the mean

D

Criterion contamination

Test Your Knowledge

In validating a newly developed adolescent resilience scale, researchers administer the new instrument alongside an established behavioral resilience rating scale completed by teachers (yielding a correlation of r = 0.65) and a standardized math computation test (yielding a correlation of r = 0.05). In the Campbell and Fiske Multitrait-Multimethod Matrix (MTMM) framework, which two forms of construct validity evidence are illustrated by these respective findings?

A

Content validity and face validity

B

Concurrent validity and predictive validity

C

Convergent validity and discriminant validity

D

Test-retest stability and scorer equivalence

Test Your Knowledge

Which of the following evaluation procedures exemplifies a Criterion-Referenced Testing (CRT) framework rather than a Norm-Referenced Testing (NRT) framework?

A

Determining whether an examinee achieved a General Weighted Average of at least 75% on the Guidance Counselor Licensure Examination to earn professional licensure

B

Reporting that a high school student scored at the 85th percentile on a nationwide scholastic aptitude assessment

C

Assigning a Stanine score of 6 to an examinee's performance on an achievement testing battery

D

Interpreting an individual's Wechsler Intelligence Scale score as a Deviation IQ of 115

Sections you finish are checked off in the contents.