13.4 Validity, Reliability, and Measurement Basics for Program Planning
Key Takeaways
Reliability is consistency (test-retest, alternate forms, internal consistency, inter-rater); validity is whether evidence supports using a score for its purpose.
A test can be reliable without being valid, but it cannot be valid without being reliable.
Pre/post items must match the lesson's objectives and selected ASCA standard (content validity); ASCA suggests three to five items per measure.
Stanines run from 1 to 9 with a mean of 5 and a standard deviation of 2; about 68% of scores fall within one standard deviation of the mean.
A drop in chronic absenteeism from 20% to 15% is a 5 percentage-point decrease and a 25% relative decrease.
13.4 Validity, Reliability, and Measurement Basics for Program Planning
Quick Summary: The Manage domain asks you to know validity and reliability as applied to program planning and to be familiar with basic measurement principles—trends, stanines, percentile ranks, validity, and reliability. Counselors use these ideas when choosing career and achievement instruments, writing pre/post assessments and needs-assessment surveys, and reading school data.
Reliability: Consistency
Reliability is the consistency of scores: a reliable measure gives similar results under similar conditions.
| Type | Question It Answers | School Example |
|---|---|---|
| Test-retest | Are scores stable over time? | An interest inventory given twice, three weeks apart |
| Alternate (parallel) forms | Do two versions give similar results? | Form A as the pre-test and Form B as the post-test |
| Internal consistency (split-half, Cronbach's alpha, KR-20) | Do the items measure the same thing? | Five items on a school-belonging scale |
| Inter-rater | Do different scorers agree? | Two counselors scoring role-plays with the same rubric |
Reliability coefficients range from 0 to 1; higher is better, and decisions about individual students call for higher reliability than evaluations of a whole program. The standard error of measurement turns reliability into a score band: SEM = SD × √(1 − r). For example, ETS reports an SEM of 5.1 for the Praxis 5422.
Validity: Accuracy of Interpretation
Validity is the degree to which evidence supports using a score for a particular purpose—in plain terms, whether a test measures what it is intended to measure. That is exactly what the ETS sample describes when a counselor wants a career assessment that "actually measures what it intends to measure."
| Type of Evidence | Meaning | Example |
|---|---|---|
| Content | Items represent the domain taught or measured | A post-test on graduation requirements covers credits, GPA, and required courses—not unrelated career facts |
| Criterion-related (concurrent or predictive) | Scores relate to an outside measure now or later | PSAT scores predicting AP exam success |
| Construct (convergent and discriminant) | Scores behave as the theory predicts | An anxiety scale correlates with other anxiety measures but not with math ability |
| Face validity | The test looks appropriate to takers (not true evidence) | Items obviously about study habits |
Key relationships: a test can be reliable without being valid (a scale that is consistently five pounds off), but it cannot be valid without being reliable. Validity is specific to a purpose and a population.
Applying Validity and Reliability to Program Planning
- Choosing instruments: check the manual for reliability, validity evidence, and a norm group that represents your students (ASCA A.14).
- Writing pre/post assessments: match every item to the learning objective and the selected ASCA standard (content validity); keep measures short—ASCA suggests three to five questions or prompts; give the same items before and after; and administer them the same way each time.
- Needs assessments: surveys of students, staff, and families identify priorities; the ETS study companion names a needs assessment as the best additional source when a counselor plans next year's academic supports after reviewing school data. Use clear, neutral wording; avoid double-barreled and leading questions; reach every group (translate for families); and follow PPRA rules when surveys touch protected topics.
- Program decisions: do not overreact to a single data point—look for consistent evidence across sources and years.
Basic Statistics for Reading School Data
- Central tendency: the mean (average) is pulled by extreme scores; the median (middle score) is better for skewed data; the mode is the most frequent score.
- Variability: the range and the standard deviation (SD).
- Normal curve: about 68% of scores fall within ±1 SD of the mean, about 95% within ±2 SD, and about 99.7% within ±3 SD.
- Percentile rank: the percentage of the norm group scoring at or below a score. The intervals are unequal, so small raw-score changes near the middle move percentiles a lot.
- Stanines: nine bands with a mean of 5 and an SD of 2; 1–3 are below average, 4–6 average, and 7–9 above average.
- Correlation (r): ranges from −1 to +1 and shows the strength and direction of a relationship, not causation.
Trends, Percent Change, and Percentage Points
Three or more years of data show whether a change is real or random fluctuation. Know the difference between two common ways of reporting change:
- Chronic absenteeism falls from 20% to 15%: that is a 5 percentage-point decrease and a 25% relative decrease (5 ÷ 20).
- In the ETS goal example, failing grades fall from 82 to 41: a 50% decrease.
Report the number of students along with percentages, and compare with a baseline or a comparison group before claiming the program caused the change.
Praxis Application
Scenario: A counselor writes a ten-question post-test for a conflict-resolution unit, but four questions ask about college admissions.
Analysis: The test has a content-validity problem: those items do not match the unit's objectives or its selected standard. The counselor should replace them with items that measure the conflict-resolution knowledge and skills the lessons taught.
A counselor gives the same career interest inventory to a group of ninth graders twice, three weeks apart, to see whether their scores stay about the same. Which property is the counselor checking?
Content validity
Predictive validity
Inter-rater reliability
Test-retest reliability
The percentage of chronically absent ninth graders fell from 20% last year to 15% this year. Which statement describes this change accurately?
A 5% relative decrease
A 5 percentage-point decrease, which is a 25% relative decrease
A 25 percentage-point decrease
A 75% relative decrease
A counselor's post-test for a lesson on test-taking strategies includes several items about career clusters. What is the main problem with this assessment?
It lacks content validity because some items do not match the lesson's learning objectives
It has too much internal consistency
It cannot be reliable because it is a post-test
It lacks face validity because students can tell what the items measure
Sections you finish are checked off in the contents.