6.6 Validity, Reliability, Objectivity & Measurement Error in CSEP-PATH Assessments
Key Takeaways
Validity is the degree to which a test measures what it claims to measure; reliability is the consistency of results when the test is repeated under the same conditions.
Objectivity, or inter-rater reliability, is the agreement between different testers measuring the same client.
The mCAFT predicts VO2max with a correlation of about 0.88 against measured treadmill VO2max (Weller et al., 1995), but individual predictions still carry an error band.
Measurement sensitivity is the smallest change a test can detect; a retest difference smaller than the test's typical error should not be presented as a real change.
Standardized CSEP-PATH procedures control errors from the client, the tester, the equipment and the environment.
6.6 Validity, Reliability, Objectivity & Measurement Error in CSEP-PATH Assessments
Three CSEP-CPT competencies deal with the quality of measurements:
- 3.25: describe the validity and reliability of the protocols in CSEP-PATH;
- 3.26: explain the measurement sensitivity of all CSEP-PATH measures;
- 3.27: identify and discuss sources of measurement error as they relate to reliability, validity and objectivity, and why it is important to minimize them.
Exam questions in this area are often conceptual ("which type of reliability?") or practical ("which error explains this result?").
Key Definitions
| Term | Meaning | CSEP-PATH example |
|---|---|---|
| Validity | The degree to which a test measures what it claims to measure | Does the mCAFT prediction match laboratory-measured VO2max? |
| Criterion validity | Agreement with a "gold standard" measure | mCAFT predicted VO2max vs. measured treadmill VO2max (r about 0.88; Weller et al., 1995) |
| Content or logical validity | The test obviously samples the quality of interest | Push-ups require upper-body muscular endurance |
| Construct validity | The test behaves as theory predicts | Trained clients score higher than untrained clients |
| Reliability | Consistency of repeated measurements under the same conditions | Two waist measurements within 0.5 cm |
| Test-retest reliability | Same client, same test, different occasions | Sit-and-reach on two days a week apart |
| Intra-rater reliability | One tester's consistency | The same trainer measuring waist circumference twice |
| Objectivity (inter-rater reliability) | Agreement between different testers | Two trainers timing the same back extension hold |
| Measurement sensitivity | The smallest change a test can detect | Is a 1 cm change in sit-and-reach real, or measurement noise? |
A test can be reliable but not valid, giving the same wrong answer every time; a bathroom scale that always reads 2 kg high is an example. A test cannot be valid unless it is reliable.
Validity and Error in Prediction Equations
Many CSEP-PATH results are predictions:
- VO2max from the mCAFT, treadmill walk, one-mile walk or cycle ergometer;
- leg power from the Sayers equation;
- 1-RM from a submaximal set.
Even a valid equation has a standard error of estimate (SEE). Individuals scatter around the prediction line, so a single client's true value may differ noticeably from the predicted one. Typical reasons:
- true maximal heart rate differs from the age-predicted value (about 10 bpm standard error);
- mechanical efficiency or stepping skill differs from the validation sample;
- the client is less like the people in the validation study (for example, outside its age range).
The PASB-Q was examined for validity and reliability against objective measures in adults (Fowles et al., 2017). Self-report tools are still affected by recall and social-desirability bias, which is why CSEP-PATH suggests logging activity for a week before completing it.
Measurement Sensitivity: When Is a Change Real?
Every measurement carries some random variation, called typical or technical error. A retest difference smaller than that error cannot be called a real change. Practical rules:
- Know the recording precision of each protocol: height to 0.5 cm, mass to 0.1 kg, waist and sit-and-reach to 0.5 cm, grip to the nearest kilogram.
- Use the protocol's own repeat rules, such as a third waist measurement when two differ by more than 0.5 cm.
- Report changes that clearly exceed day-to-day variation, and confirm surprising changes with a repeat test.
- Health Benefit Rating bands are wide. A client may improve meaningfully without changing band, or change band with a small, possibly unreal, difference near a boundary.
Sources of Measurement Error and How CSEP-PATH Controls Them
| Source | Examples | Control |
|---|---|---|
| Client | Caffeine, recent meal or vigorous exercise; poor sleep; illness; anxiety; motivation; unfamiliarity with the task | Welcome Letter pre-test instructions; quiet rest; clear demonstration and practice; consistent encouragement |
| Tester | Wrong landmark; inconsistent cueing; mistimed pulse counts; reading errors; bias | Follow CSEP-PATH protocols exactly; train and practise; use the same tester for retests |
| Equipment | Uncalibrated sphygmomanometer, dynamometer or ergometer; wrong cuff size; worn measuring tape | Calibration and maintenance schedules; correct cuff sizing; check zero before use |
| Environment | Heat, humidity, noise, distractions, time of day | Controlled room temperature; quiet space; test at a similar time of day |
| Protocol | Changing order of tests, rest periods, warm-up or footwear between visits | Use the same order, warm-up, clothing and footwear each time |
Why it matters: measurement error can lead to wrong safety decisions, such as missing an elevated blood pressure. It can produce misleading Health Benefit Ratings, and it can show false progress or false decline that hurts client motivation and trust. Accurate measurement is also part of the standard of care a court would expect from a qualified exercise professional.
Two trainers independently time the same client's back extension test and record 92 seconds and 104 seconds. Which measurement property does this disagreement mainly reflect?
Criterion validity against a laboratory standard
Construct validity of the endurance test
Measurement sensitivity of the stopwatch
Objectivity (inter-rater reliability)
A bathroom-style scale consistently reads 2.0 kg higher than a calibrated medical scale every time the same client steps on it. How should this scale be described?
Valid but not reliable
Reliable but not valid
Neither reliable nor sensitive
Both valid and reliable
A client's sit-and-reach improves from 31.0 cm to 31.5 cm after four weeks. What is the most appropriate interpretation?
A meaningful improvement, because any recorded increase in reach must be a real change.
A decline, because genuine flexibility gains over four weeks are usually much larger.
A change too small to call real: it equals the 0.5 cm recording precision of the protocol.
Proof that the sit-and-reach test is invalid for this client and should be replaced.
Sections you finish are checked off in the contents.