5.2 Content, Criterion-Related (Predictive/Concurrent), and Construct Validity in Public Selection

Key Takeaways

  • Validity is the degree to which accumulated empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and decisions made based on test scores.
  • Content validity demonstrates that test items represent a direct, representative sample of critical job behaviors, duties, or knowledge/skills identified through a comprehensive job analysis, and is quantified via Lawshe's Content Validity Ratio (CVR).
  • Criterion-related validity establishes a statistical correlation (validity coefficient r) between test scores (predictor) and actual job performance metrics (criterion), utilizing either predictive (prospective) or concurrent (incumbent) research designs.
  • Construct validity proves that an assessment instrument measures an underlying psychological construct (e.g., cognitive ability, spatial reasoning, emotional intelligence) and that this construct is critical for successful job performance.
  • UGESP explicitly prohibits relying on content validity for tests measuring abstract mental traits, psychological constructs, or knowledge/skills that can be easily acquired during brief post-hire orientation.
Last updated: August 2026

5.2 Content, Criterion-Related (Predictive/Concurrent), and Construct Validity in Public Selection

In public personnel management, test validation is the scientific and legal process of accumulating empirical evidence to demonstrate that inferences, rankings, and employment decisions drawn from assessment scores are justified, accurate, and job-related. Under Section 14 of the Uniform Guidelines on Employee Selection Procedures (UGESP), public employers defending a selection procedure that produces adverse impact must establish validity through one of three foundational validation strategies:

  1. Content Validity
  2. Criterion-Related Validity (Predictive or Concurrent)
  3. Construct Validity

Each validation model possesses distinct methodological requirements, statistical benchmarks, practical advantages, and legal limitations within civil service environments.


1. The Tripartite Validation Framework

+-----------------------------------------------------------------------------------------+
|                    THE THREE CLASSICAL VALIDATION STRATEGIES (UGESP)                    |
|                                                                                         |
|   +------------------------------------+   +------------------------------------+       |
|   |         CONTENT VALIDITY           |   |     CRITERION-RELATED VALIDITY     |       |
|   | - Direct behavioral sample of job  |   | - Statistical correlation between  |       |
|   | - Based on thorough job analysis   |   |   test score (X) & performance (Y) |       |
|   | - Work samples, knowledge tests    |   | - Predictive or Concurrent designs |       |
|   | - Quantified via Lawshe's CVR      |   | - Evaluated by correlation coeff r |       |
|   +------------------------------------+   +------------------------------------+       |
|                     \                                         /                         |
|                      \                                       /                          |
|                       v                                     v                           |
|                     +-----------------------------------------+                         |
|                     |           CONSTRUCT VALIDITY            |                         |
|                     | - Measures underlying psychological     |                         |
|                     |   trait/construct (GMA, integrity)      |                         |
|                     | - Requires convergent/discriminant proof|                         |
|                     | - Most complex and demanding validation |                         |
|                     +-----------------------------------------+                         |
+-----------------------------------------------------------------------------------------+

2. Content Validity: Principles and Public Sector Application

Content validity demonstrates that the content of the selection instrument is a representative sample of the critical work behaviors, operational duties, or prerequisite Knowledge, Skills, and Abilities (KSAs) required for effective performance in the target position.

Core Prerequisites for Content Validity under UGESP Section 14C:

  1. Rigorous Job Analysis: Must identify critical job tasks, their frequency and importance, and the specific KSAs required to execute each task.
  2. Direct Operational Linkage: The test content must closely replicate the actual work environment, duties, or knowledge domains (e.g., a typing test for an administrative assistant, a drafting simulation for an engineer, a physical agility test for a firefighter).
  3. Pre-Hire Prerequisite Rule: Content validity cannot be used to assess knowledge, skills, or abilities that are routinely taught during post-hire onboarding or learned on the job within a brief orientation period.
  4. The Abstract Trait Prohibition: UGESP explicitly states that content validity is inappropriate for tests measuring abstract mental traits, psychological constructs, or general aptitudes (e.g., intelligence, judgment, leadership orientation, neuroticism).
+-----------------------------------------------------------------------------------------+
|                  LAWSHE'S CONTENT VALIDITY RATIO (CVR) METHODOLOGY                      |
|                                                                                         |
|   A panel of Subject Matter Experts (SMEs) evaluates each proposed test item using      |
|   a 3-point scale:                                                                      |
|   1. Essential to job performance                                                       |
|   2. Useful but not essential                                                           |
|   3. Not necessary                                                                      |
|                                                                                         |
|                           n_e - (N / 2)                                                 |
|                  CVR  =  ---------------                                                |
|                              N / 2                                                      |
|                                                                                         |
|   Where:                                                                                |
|   n_e = Number of SME panelists rating the item as "Essential"                          |
|   N   = Total number of SME panelists on the review board                               |
+-----------------------------------------------------------------------------------------+

Interpreting Lawshe's CVR:

  • CVR Range: Ranges from -1.00 (all SMEs rate item "Not Necessary") to 0.00 (exactly half rate "Essential") to +1.00 (all SMEs rate "Essential").
  • Statistical Significance Threshold: An item is retained only if its CVR exceeds the minimum statistical threshold based on panel size ($p < .05$, one-tailed test):
SME Panel Size ($N$)Minimum Required CVR Value ($p < .05$)
5 SMEs0.99 (Unanimous agreement required)
10 SMEs0.62
15 SMEs0.49
20 SMEs0.42
30 SMEs0.33
40+ SMEs0.29

Content Validity Index (CVI)

The Content Validity Index (CVI) represents the arithmetic mean of all retained item CVR values across the final assessment instrument, providing an overall index of the test's content representativeness.


3. Criterion-Related Validity: Predictive vs. Concurrent Designs

Criterion-related validity demonstrates a statistically significant empirical relationship between scores on an assessment instrument (the predictor, $X$) and subsequent measures of actual work performance (the criterion, $Y$).

+-----------------------------------------------------------------------------------------+
|                 CRITERION-RELATED VALIDITY: RESEARCH DESIGN COMPARISON                  |
|                                                                                         |
|   [PREDICTIVE VALIDITY DESIGN (PROSPECTIVE)]                                            |
|                                                                                         |
|   Step 1: Administer Test (X) to All APPLICANTS at Initial Screening                   |
|   Step 2: Hire Applicants WITHOUT Using Test Scores (or Hire All Applicants)            |
|   Step 3: Allow Employees to Work in Target Classification (6-12 Months)                |
|   Step 4: Collect Objective Criterion Performance Data (Y) (Ratings, Errors, Output)   |
|   Step 5: Compute Correlation Coefficient r_xy between Applicant Score & Performance    |
|                                                                                         |
|   -----------------------------------------------------------------------------------   |
|                                                                                         |
|   [CONCURRENT VALIDITY DESIGN (INCUMBENT)]                                              |
|                                                                                         |
|   Step 1: Select Representative Sample of CURRENT TENURED EMPLOYEES                     |
|   Step 2: Administer Test (X) to Current Incumbents                                     |
|   Step 3: Collect SIMULTANEOUS Performance Evaluations (Y) for Same Incumbents         |
|   Step 4: Compute Correlation Coefficient r_xy between Incumbent Score & Performance    |
+-----------------------------------------------------------------------------------------+

Comprehensive Comparison Matrix: Predictive vs. Concurrent

Evaluation DimensionPredictive Validity DesignConcurrent Validity Design
Test PopulationJob Applicants (pre-hire).Current Job Incumbents (post-hire).
Timing of DataPredictor collected at baseline; criterion collected months later.Predictor and criterion collected simultaneously.
Range RestrictionMinimal. Evaluates full spectrum of applicant abilities.Severe. Poor performers have been fired; top performers promoted.
Test-Taker MotivationHigh (applicants striving to secure employment).Variable/Low (incumbents may experience test fatigue or anxiety).
Organizational Cost & TimeHigh cost; significant time lag before validation results emerge.Fast turnaround; immediate empirical results.
Legal/Scientific RobustnessConsidered the gold standard of criterion validation.Legally acceptable under UGESP, but subject to range restriction critiques.

Statistical Evaluation of Validity Coefficients ($r_{xy}$)

Criterion validity is quantified using the Pearson product-moment correlation coefficient ($r_{xy}$) between the predictor and criterion:

rxy=(XXˉ)(YYˉ)(XXˉ)2(YYˉ)2r_{xy} = \frac{\sum (X - \bar{X})(Y - \bar{Y})}{\sqrt{\sum (X - \bar{X})^2 \sum (Y - \bar{Y})^2}}

  • Statistical Significance: UGESP requires that the correlation coefficient be statistically significant at $p < .05$ (meaning there is less than a 5% probability that the correlation occurred by random sampling error).
  • Coefficient of Determination ($r^2$): Represents the proportion of variance in job performance explained by the test score (e.g., if $r = .50$, $r^2 = .25$, meaning the test explains 25% of the performance variance).

Professional Benchmarks for Validity Coefficients ($r$):

Validity Coefficient ($r$)Practical Interpretation in Employee Selection
$r < .10$Negligible / Unusable for selection decisions.
$r = .11 - .20$Weak / Minimally useful as a standalone hurdle.
$r = .21 - .34$Acceptable / Useful when combined in a composite battery.
$r = .35 - .50$Strong / Highly effective predictor of performance.
$r > .50$Exceptional / Rare in real-world selection systems.

Range Restriction (Attenuation) and Correction Formulas

When calculating validity coefficients on current employees (concurrent validity) or pre-screened applicants, the observed correlation is artificially depressed because low-scoring individuals are absent from the sample. Industrial psychologists apply Thorndike's Case II Range Restriction Correction Formula to estimate the true operational validity coefficient ($r_U$) in the unrestricted applicant pool:

rU=rR(SDUSDR)1+rR2((SDUSDR)21)r_U = \frac{r_R \left(\frac{SD_U}{SD_R}\right)}{\sqrt{1 + r_R^2 \left(\left(\frac{SD_U}{SD_R}\right)^2 - 1\right)}}

Where $r_R$ is the restricted sample correlation, $SD_R$ is the standard deviation in the restricted sample, and $SD_U$ is the standard deviation in the unrestricted applicant pool.


4. Criterion Measurement: Contamination, Deficiency, and Reliability

A criterion validation study is only as sound as the job performance measure (criterion) it utilizes. UGESP Section 14B(3) requires that criteria represent critical aspects of work performance.

+-----------------------------------------------------------------------------------------+
|                        CRITERION DYNAMICS IN VALIDATION STUDIES                         |
|                                                                                         |
|       +-------------------------------------------------------------------------+       |
|       |                      THE ACTUAL WORK DOMAIN                             |       |
|       |       +---------------------------------+                               |       |
|       |       |       CRITERION DEFICIENCY      |                               |       |
|       |       | (Critical job facets omitted)   |                               |       |
|       |       | - Customer de-escalation skill  |                               |       |
|       |       | - Team safety adherence         |                               |       |
|       |       +---------------------------------+                               |       |
|       |                 |                                                       |       |
|       |                 v                                                       |       |
|       |       +=================================+ <--- [VALID CRITERION OVERLAP]|       |
|       |       |       CRITERION RELEVANCE       |      (Accurate measure of     |       |
|       |       | (Core job duties captured)      |       true job competence)    |       |
|       |       +=================================+                               |       |
|       +-----------------|-------------------------------------------------------+       |
|                         |                                                               |
|                         v                                                               |
|       +-----------------------------------------+                                       |
|       |         CRITERION CONTAMINATION         |                                       |
|       | (Non-job factors polluting the metric)  |                                       |
|       | - Rater personal favoritism             |                                       |
|       | - Shift equipment differences           |                                       |
|       | - Unequal territory sales opportunities |                                       |
|       +-----------------------------------------+                                       |
+-----------------------------------------------------------------------------------------+
  • Criterion Relevance: The degree to which the criterion measures the essential performance dimensions of the job.
  • Criterion Deficiency: Failure of the criterion to measure important job duties (e.g., evaluating a police officer exclusively on the number of traffic citations issued, omitting community problem-solving and de-escalation).
  • Criterion Contamination: Inclusion of extraneous, non-job-related variance in the criterion score (e.g., subjective supervisory ratings influenced by employee personality, race, gender, or differences in mechanical equipment quality across shifts).

5. Construct Validity: Psychological Constructs in Selection

Construct validity demonstrates that an assessment instrument measures an underlying psychological construct, trait, or characteristic (e.g., general mental ability, mechanical comprehension, conscientiousness, emotional intelligence) AND that this construct is critical for successful job performance.

UGESP Section 14D Stringency

Construct validity is the most intellectually rigorous and legally complex validation strategy. To establish construct validity under UGESP, the agency must provide a three-part chain of evidence:

  1. Demonstrate the Construct Exists: Theoretical and empirical proof that the construct represents a distinct, coherent psychological domain.
  2. Demonstrate Test Measures the Construct: Evidence of convergent validity (high correlation with established tests measuring the same construct) and discriminant validity (low correlation with tests measuring unrelated constructs).
  3. Demonstrate Construct Predicts Job Performance: Empirical criterion-related evidence linking the construct to successful performance across the target job family.

6. Validity Generalization (VG) and Synthetic Validation

In modern public HR administration, small agencies frequently lack the large sample sizes required to conduct localized, highly powered criterion-related validation studies ($N \ge 100$).

Validity Generalization (Hunter-Schmidt Model)

Meta-analytic research demonstrates that validities for well-defined cognitive and behavioral predictors do not vary randomly from organization to organization (the "situational specificity hypothesis" is largely a statistical artifact of small sample error and range restriction). Under professional SIOP standards, an agency can transport validity evidence from broad meta-analyses if the target job is demonstrably similar in essential work duties and required KSAOs.

UGESP Transportability Standards (Section 7)

To legally transport a validity study from another jurisdiction under UGESP, the receiving public agency must prove:

  1. Job Comparability: Thorough job analyses demonstrate that the essential duties and critical work behaviors of the two positions are substantially identical.
  2. Test Comparability: The identical assessment instrument is utilized under standardized administration conditions.
  3. Applicant Pool Comparability: The demographic and educational characteristics of the applicant pools are substantially similar.
Loading diagram...
Methodological Architecture of the Three Validation Models
Test Your Knowledge

An HR analyst is utilizing Lawshe's methodology to validate a 50-item written exam for Fire Captain. If a review panel comprises 20 Subject Matter Experts (SMEs), and exactly 16 SMEs rate a specific item as 'Essential', what is the Content Validity Ratio (CVR) for that item?

A
B
C
D
Test Your Knowledge

Under UGESP Section 14C, which of the following assessment types is EXPLICITLY prohibited from being validated using a content validity strategy alone?

A
B
C
D
Test Your Knowledge

In employee selection research, what is the primary psychometric limitation associated with a concurrent criterion-related validity design compared to a predictive design?

A
B
C
D
Test Your Knowledge

When validating a selection system for municipal water treatment plant operators, an HR team discovers that supervisory performance evaluations are heavily biased by whether an operator works the day or night shift due to older filtration machinery on the night shift. In psychometric terms, this performance criterion suffers from:

A
B
C
D