13.4 Validity, Reliability, and Measurement Basics for Program Planning

Key Takeaways

  • Reliability is consistency (test-retest, alternate forms, internal consistency, inter-rater); validity is whether evidence supports using a score for its purpose.

  • A test can be reliable without being valid, but it cannot be valid without being reliable.

  • Pre/post items must match the lesson's objectives and selected ASCA standard (content validity); ASCA suggests three to five items per measure.

  • Stanines run from 1 to 9 with a mean of 5 and a standard deviation of 2; about 68% of scores fall within one standard deviation of the mean.

  • A drop in chronic absenteeism from 20% to 15% is a 5 percentage-point decrease and a 25% relative decrease.

Last updated: September 2026

13.4 Validity, Reliability, and Measurement Basics for Program Planning

Quick Summary: The Manage domain asks you to know validity and reliability as applied to program planning and to be familiar with basic measurement principles—trends, stanines, percentile ranks, validity, and reliability. Counselors use these ideas when choosing career and achievement instruments, writing pre/post assessments and needs-assessment surveys, and reading school data.


Reliability: Consistency

Reliability is the consistency of scores: a reliable measure gives similar results under similar conditions.

TypeQuestion It AnswersSchool Example
Test-retestAre scores stable over time?An interest inventory given twice, three weeks apart
Alternate (parallel) formsDo two versions give similar results?Form A as the pre-test and Form B as the post-test
Internal consistency (split-half, Cronbach's alpha, KR-20)Do the items measure the same thing?Five items on a school-belonging scale
Inter-raterDo different scorers agree?Two counselors scoring role-plays with the same rubric

Reliability coefficients range from 0 to 1; higher is better, and decisions about individual students call for higher reliability than evaluations of a whole program. The standard error of measurement turns reliability into a score band: SEM = SD × √(1 − r). For example, ETS reports an SEM of 5.1 for the Praxis 5422.

Validity: Accuracy of Interpretation

Validity is the degree to which evidence supports using a score for a particular purpose—in plain terms, whether a test measures what it is intended to measure. That is exactly what the ETS sample describes when a counselor wants a career assessment that "actually measures what it intends to measure."

Type of EvidenceMeaningExample
ContentItems represent the domain taught or measuredA post-test on graduation requirements covers credits, GPA, and required courses—not unrelated career facts
Criterion-related (concurrent or predictive)Scores relate to an outside measure now or laterPSAT scores predicting AP exam success
Construct (convergent and discriminant)Scores behave as the theory predictsAn anxiety scale correlates with other anxiety measures but not with math ability
Face validityThe test looks appropriate to takers (not true evidence)Items obviously about study habits

Key relationships: a test can be reliable without being valid (a scale that is consistently five pounds off), but it cannot be valid without being reliable. Validity is specific to a purpose and a population.

Applying Validity and Reliability to Program Planning

  1. Choosing instruments: check the manual for reliability, validity evidence, and a norm group that represents your students (ASCA A.14).
  2. Writing pre/post assessments: match every item to the learning objective and the selected ASCA standard (content validity); keep measures short—ASCA suggests three to five questions or prompts; give the same items before and after; and administer them the same way each time.
  3. Needs assessments: surveys of students, staff, and families identify priorities; the ETS study companion names a needs assessment as the best additional source when a counselor plans next year's academic supports after reviewing school data. Use clear, neutral wording; avoid double-barreled and leading questions; reach every group (translate for families); and follow PPRA rules when surveys touch protected topics.
  4. Program decisions: do not overreact to a single data point—look for consistent evidence across sources and years.

Basic Statistics for Reading School Data

  • Central tendency: the mean (average) is pulled by extreme scores; the median (middle score) is better for skewed data; the mode is the most frequent score.
  • Variability: the range and the standard deviation (SD).
  • Normal curve: about 68% of scores fall within ±1 SD of the mean, about 95% within ±2 SD, and about 99.7% within ±3 SD.
  • Percentile rank: the percentage of the norm group scoring at or below a score. The intervals are unequal, so small raw-score changes near the middle move percentiles a lot.
  • Stanines: nine bands with a mean of 5 and an SD of 2; 1–3 are below average, 4–6 average, and 7–9 above average.
  • Correlation (r): ranges from −1 to +1 and shows the strength and direction of a relationship, not causation.

Trends, Percent Change, and Percentage Points

Three or more years of data show whether a change is real or random fluctuation. Know the difference between two common ways of reporting change:

  • Chronic absenteeism falls from 20% to 15%: that is a 5 percentage-point decrease and a 25% relative decrease (5 ÷ 20).
  • In the ETS goal example, failing grades fall from 82 to 41: a 50% decrease.

Report the number of students along with percentages, and compare with a baseline or a comparison group before claiming the program caused the change.

Praxis Application

Scenario: A counselor writes a ten-question post-test for a conflict-resolution unit, but four questions ask about college admissions.

Analysis: The test has a content-validity problem: those items do not match the unit's objectives or its selected standard. The counselor should replace them with items that measure the conflict-resolution knowledge and skills the lessons taught.

Loading diagram...
Measurement Quality Checks for Program Planning
Test Your Knowledge

A counselor gives the same career interest inventory to a group of ninth graders twice, three weeks apart, to see whether their scores stay about the same. Which property is the counselor checking?

A

Content validity

B

Predictive validity

C

Inter-rater reliability

D

Test-retest reliability

Test Your Knowledge

The percentage of chronically absent ninth graders fell from 20% last year to 15% this year. Which statement describes this change accurately?

A

A 5% relative decrease

B

A 5 percentage-point decrease, which is a 25% relative decrease

C

A 25 percentage-point decrease

D

A 75% relative decrease

Test Your Knowledge

A counselor's post-test for a lesson on test-taking strategies includes several items about career clusters. What is the main problem with this assessment?

A

It lacks content validity because some items do not match the lesson's learning objectives

B

It has too much internal consistency

C

It cannot be reliable because it is a post-test

D

It lacks face validity because students can tell what the items measure

Sections you finish are checked off in the contents.