15.4 Assessment Types, Methods, Validity, and Reliability
Key Takeaways
- Formative assessment occurs during learning to adjust instruction, summative assessment certifies achievement afterward, and diagnostic assessment establishes baseline before instruction begins.
- Criterion-referenced assessment compares a student to a defined standard that every student can meet, while norm-referenced assessment compares students to one another so a fixed share must fall below any percentile.
- Validity is the degree to which an assessment measures what it claims to measure, and the dominant validity failure in physical education is misalignment -- assessing a demonstration standard with a recall test.
- Reliability is necessary but not sufficient for validity -- a stopwatch measures a mile time reliably without validly measuring whether the student learned pacing strategy -- and inter-rater reliability is the practical concern in observational physical education, improved by specific observable criteria, rater training, and calibration against a common performance.
- Grading in physical education must be standards-based, reflecting learning in the psychomotor, cognitive, and affective domains rather than non-academic criteria such as dressing out or attendance.
Student Assessment Strategies
Assessment in physical education is the systematic process of gathering, analyzing, and interpreting evidence of student learning to make informed instructional decisions and evaluate student progress. Far from being a mere end-of-unit grading ritual, high-quality assessment drives student motivation, informs differentiated instruction, and provides accountability for physical literacy. The Praxis Health and Physical Education (5857) exam assesses expertise in formative versus summative assessment, criterion-referenced versus norm-referenced evaluation, the FitnessGram fitness battery, rubric construction, motor skill checklists, student self-assessment, and standards-based grading practices.
1. Formative vs. Summative & Assessment Frameworks
Physical educators utilize two complementary forms of assessment to monitor and evaluate student growth:
Formative vs. Summative Assessment
- Formative Assessment (Assessment FOR Learning): Conducted continuously during the instructional process. Purpose: to diagnose learning gaps, provide immediate corrective feedback to students, and allow teachers to adjust daily instruction. Examples include teacher check-ins, thumb polls, quick skill checklists during warm-ups, peer observation cards, and exit tickets.
- Summative Assessment (Assessment OF Learning): Conducted at the conclusion of a unit or term. Purpose: to evaluate cumulative student achievement against standards-based learning objectives. Examples include end-of-unit skill rubrics, written tactical exams, portfolio submissions, and formal fitness post-tests.
Criterion-Referenced vs. Norm-Referenced Evaluation
| Assessment Type | Evaluation Standard | Purpose & Example in PE |
|---|---|---|
| Criterion-Referenced | Student performance is compared against an absolute, predetermined standard or criterion of mastery. | FitnessGram Healthy Fitness Zones (HFZ): A student's score is evaluated against health-related standards established by medical research, regardless of how peers perform. |
| Norm-Referenced | Student performance is compared against the performance of a normative peer group (percentiles). | President's Challenge Percentiles: Comparing a student's mile run time to a national average (e.g., '85th percentile'). Note: Deprecated in modern standards-based PE. |
Modern physical education strongly emphasizes criterion-referenced assessment because it promotes individual growth and health-related standards rather than discouraging less athletic students through peer comparison.
2. The Full Range of Assessment Types and Methods
The ETS blueprint names formative, summative, authentic, portfolio, standardized, rubric, criterion-referenced, and norm-referenced assessment. They are not a single list -- they answer different questions.
| Type | What it is | Physical education example |
|---|---|---|
| Formative | Assessment during learning, used to adjust instruction | Observed practice trial with immediate corrective feedback |
| Summative | Assessment after learning, certifying achievement | End-of-unit skill demonstration scored on a rubric |
| Diagnostic | Assessment before instruction, establishing baseline | Preassessment of a striking pattern in week one |
| Authentic | A task resembling a genuine real-world application | Officiating a live game; designing and following a personal fitness plan |
| Portfolio | A selected, reflected-on collection of work over time | Video of a skill at three points in the unit with the student's written analysis |
| Standardized | Administered and scored under uniform conditions | FITNESSGRAM administered per protocol |
| Rubric-scored | Performance judged against described criteria and levels | Analytic rubric for a dance composition |
| Criterion-referenced | Performance compared to a defined standard | FITNESSGRAM Healthy Fitness Zone; a skill rubric's proficient level |
| Norm-referenced | Performance compared to other people | A percentile rank; the Presidential Physical Fitness Award's 85th-percentile benchmark |
Criterion- versus norm-referenced is the distinction most often tested. Criterion-referenced assessment asks whether the student met the standard, and every student can meet it. Norm-referenced assessment asks how the student ranks, and by construction a fixed share must fall below any given percentile. Instructional assessment in physical education should be criterion-referenced; norm-referenced comparison is appropriate for research and program benchmarking, not for judging an individual student's learning.
A portfolio is not a folder of everything a student produced. Selection by the student, inclusion of drafts alongside finished work to show growth, and a written reflection are what make it assessment rather than storage.
3. Validity, Reliability, Bias, and Interpreting Results
These four concepts determine whether an assessment result means anything.
Validity
Validity is the degree to which an assessment measures what it claims to measure and supports the interpretation being made from it.
| Type | Question | Physical education example |
|---|---|---|
| Content validity | Does the assessment sample the content and skills it claims to cover? | A volleyball assessment that includes serving, passing, and setting rather than only serving |
| Construct validity | Does it measure the underlying construct? | Does a written test on refusal techniques actually measure refusal skill? It does not |
| Criterion validity | Does it agree with an established measure? | Does the PACER estimate correspond to laboratory-measured aerobic capacity? |
| Face validity | Does it appear to measure the right thing to the test-taker? | Weakest form, but it affects student buy-in |
The dominant validity failure in physical education is misalignment: assessing a demonstration standard with a recall test, assessing a communication objective with a poster, or grading a student's fitness score as evidence of what they learned. In each case the score is reliable and precise and measures the wrong thing.
Reliability
Reliability is consistency -- would the same performance produce the same score again?
| Type | Question |
|---|---|
| Test-retest | Would the same student score similarly on a second administration under the same conditions? |
| Inter-rater | Would two trained observers score the same performance the same way? |
| Intra-rater | Would the same observer score the same performance the same way on two occasions? |
| Internal consistency | Do items intended to measure the same thing agree with one another? |
Inter-rater reliability is the practical concern in physical education, because so much assessment is observational. It is improved by specific observable criteria rather than global impressions, training observers on the criteria, calibrating raters against a common recorded performance, and limiting the number of criteria judged at once.
Reliability is necessary but not sufficient for validity. A stopwatch measures a mile time very reliably; it is not a valid measure of whether the student learned pacing strategy.
Bias
Bias is systematic error that advantages or disadvantages a group of students for reasons unrelated to what is being assessed.
| Source | How it appears | Countermeasure |
|---|---|---|
| Halo effect | A skilled or well-liked student is rated higher on effort and cooperation than an equally engaged peer | Score against defined observable criteria, one criterion at a time |
| Prior-knowledge and experience bias | Assessment rewards out-of-school sport exposure rather than in-class learning | Assess growth from the student's own baseline; assess taught content |
| Language bias | A written assessment measures reading and writing rather than the physical education construct | Provide translation, visuals, oral or demonstrated response options |
| Cultural bias | Scenario contexts, demeanor expectations, or eye-contact norms disadvantage some students | Review scenarios and criteria for cultural assumptions |
| Body and maturation bias | Items favor early-maturing or larger students | Use criterion-referenced health standards; never grade fitness scores |
| Order and fatigue effects | Students assessed last perform differently | Rotate assessment order |
Interpreting results
- Ask what the score licenses. A single observation supports a tentative conclusion, not a final judgment.
- Look at the class pattern, the spread, and the outliers, not only the average.
- Separate learning from circumstance. A low fitness score may reflect maturation and out-of-school resources rather than effort or instruction.
- Triangulate. Combine observation, a cognitive check, and student self-report before drawing a conclusion about a student.
- Report against criteria and growth, never as a rank.
3. Rubrics & Motor Skill Checklists
To assess psychomotor skill execution objectively, physical educators use rubrics and motor skill checklists.
Analytic vs. Holistic Rubrics
- Analytic Rubrics: Break a motor performance down into distinct components or critical elements (e.g., stance, arm backswing, plant foot, follow-through) and score each criterion independently on a 1-4 scale. Provides detailed diagnostic feedback.
- Holistic Rubrics: Provide a single, overall score (e.g., Advanced, Proficient, Basic, Below Basic) based on an aggregate impression of the total movement performance.
Motor Skill Checklist Example: Overhand Throw Mechanics
| Critical Element | Observed (✓) / Needs Work (✗) | Instructional Cue |
|---|---|---|
| 1. Side orientation to target | ✓ | 'Point non-throwing shoulder to target' |
| 2. Wind-up / Deep arm arc | ✓ | 'Make a big L with throwing arm' |
| 3. Step with opposite foot | ✓ | 'Step toward target with front foot' |
| 4. Hip and trunk rotation | ✗ | 'Uncoil your hips and chest' |
| 5. High release & diagonal follow-through | ✓ | 'Reach across to opposite pocket' |
4. Standards-Based Grading & Invalid Practices
Grading in physical education must reflect student achievement of educational standards across psychomotor, cognitive, and affective domains.
Invalid vs. Valid PE Grading Practices
INVALID PE GRADING PRACTICES (Avoid on Praxis Exam) VALID STANDARDS-BASED GRADING
├── Grading solely on attendance or 'dressing out' ├── Psychomotor skill rubrics
├── Grading strictly on subjective 'effort' ├── Cognitive knowledge exit tickets
├── Using fitness test scores as 100% of grade ├── FitnessGram goal-setting logs
└── Awarding extra credit for bringing gym towels └── Affective social behavior rubrics
Using FitnessGram scores directly to determine grades is considered invalid practice by SHAPE America because fitness test results are heavily influenced by genetics, body type, maturation, and out-of-school factors. Fitness tests should be used for goal setting and personal awareness, not for assigning report card grades.
Common Praxis Exam Traps & Real-World Scenarios
- Trap 1: Grading PE on Dressing Out. Exam questions frequently highlight scenarios where a teacher assigns grades based on bringing gym clothes. This is an invalid grading practice because it measures compliance/resources rather than standard-based physical literacy.
- Trap 2: Norm-Referenced Fitness Testing. Identifying student rank by comparing them to classmates (e.g., 'Johnny was the 3rd fastest runner in class') is norm-referenced. Modern PE uses criterion-referenced Healthy Fitness Zones.
- Trap 3: Misinterpreting PACER Form Breaks. In the FitnessGram PACER test, if a student fails to reach the line before the beep, they receive a warning. The test ends when the student fails to reach the line for the second time (not the first).
A physical education teacher observes students during a badminton lead-up activity, filling out a quick clipboard checklist on racquet grip and footwork to provide immediate verbal corrections. What type of assessment is being conducted?
A teacher's stated objective is that students will demonstrate appropriate pacing strategy in a distance run. The teacher assesses it by recording each student's finishing time with a stopwatch. How should this assessment be characterized?
A middle school teacher creates a psychomotor assessment for a tennis forehand stroke that evaluates four separate components—Ready Stance, Racquet Backswing, Contact Point, and Follow-Through—scoring each component on a scale from 1 to 4 with specific descriptive criteria. What type of assessment tool is this?
According to SHAPE America guidelines, which of the following grading policies represents an invalid assessment practice in physical education?