6.6 Validity, Reliability, Objectivity & Measurement Error in CSEP-PATH Assessments

Key Takeaways

  • Validity is the degree to which a test measures what it claims to measure; reliability is the consistency of results when the test is repeated under the same conditions.

  • Objectivity, or inter-rater reliability, is the agreement between different testers measuring the same client.

  • The mCAFT predicts VO2max with a correlation of about 0.88 against measured treadmill VO2max (Weller et al., 1995), but individual predictions still carry an error band.

  • Measurement sensitivity is the smallest change a test can detect; a retest difference smaller than the test's typical error should not be presented as a real change.

  • Standardized CSEP-PATH procedures control errors from the client, the tester, the equipment and the environment.

Last updated: October 2026

6.6 Validity, Reliability, Objectivity & Measurement Error in CSEP-PATH Assessments

Three CSEP-CPT competencies deal with the quality of measurements:

  • 3.25: describe the validity and reliability of the protocols in CSEP-PATH;
  • 3.26: explain the measurement sensitivity of all CSEP-PATH measures;
  • 3.27: identify and discuss sources of measurement error as they relate to reliability, validity and objectivity, and why it is important to minimize them.

Exam questions in this area are often conceptual ("which type of reliability?") or practical ("which error explains this result?").


Key Definitions

TermMeaningCSEP-PATH example
ValidityThe degree to which a test measures what it claims to measureDoes the mCAFT prediction match laboratory-measured VO2max?
Criterion validityAgreement with a "gold standard" measuremCAFT predicted VO2max vs. measured treadmill VO2max (r about 0.88; Weller et al., 1995)
Content or logical validityThe test obviously samples the quality of interestPush-ups require upper-body muscular endurance
Construct validityThe test behaves as theory predictsTrained clients score higher than untrained clients
ReliabilityConsistency of repeated measurements under the same conditionsTwo waist measurements within 0.5 cm
Test-retest reliabilitySame client, same test, different occasionsSit-and-reach on two days a week apart
Intra-rater reliabilityOne tester's consistencyThe same trainer measuring waist circumference twice
Objectivity (inter-rater reliability)Agreement between different testersTwo trainers timing the same back extension hold
Measurement sensitivityThe smallest change a test can detectIs a 1 cm change in sit-and-reach real, or measurement noise?

A test can be reliable but not valid, giving the same wrong answer every time; a bathroom scale that always reads 2 kg high is an example. A test cannot be valid unless it is reliable.


Validity and Error in Prediction Equations

Many CSEP-PATH results are predictions:

  • VO2max from the mCAFT, treadmill walk, one-mile walk or cycle ergometer;
  • leg power from the Sayers equation;
  • 1-RM from a submaximal set.

Even a valid equation has a standard error of estimate (SEE). Individuals scatter around the prediction line, so a single client's true value may differ noticeably from the predicted one. Typical reasons:

  • true maximal heart rate differs from the age-predicted value (about 10 bpm standard error);
  • mechanical efficiency or stepping skill differs from the validation sample;
  • the client is less like the people in the validation study (for example, outside its age range).

The PASB-Q was examined for validity and reliability against objective measures in adults (Fowles et al., 2017). Self-report tools are still affected by recall and social-desirability bias, which is why CSEP-PATH suggests logging activity for a week before completing it.


Measurement Sensitivity: When Is a Change Real?

Every measurement carries some random variation, called typical or technical error. A retest difference smaller than that error cannot be called a real change. Practical rules:

  • Know the recording precision of each protocol: height to 0.5 cm, mass to 0.1 kg, waist and sit-and-reach to 0.5 cm, grip to the nearest kilogram.
  • Use the protocol's own repeat rules, such as a third waist measurement when two differ by more than 0.5 cm.
  • Report changes that clearly exceed day-to-day variation, and confirm surprising changes with a repeat test.
  • Health Benefit Rating bands are wide. A client may improve meaningfully without changing band, or change band with a small, possibly unreal, difference near a boundary.

Sources of Measurement Error and How CSEP-PATH Controls Them

SourceExamplesControl
ClientCaffeine, recent meal or vigorous exercise; poor sleep; illness; anxiety; motivation; unfamiliarity with the taskWelcome Letter pre-test instructions; quiet rest; clear demonstration and practice; consistent encouragement
TesterWrong landmark; inconsistent cueing; mistimed pulse counts; reading errors; biasFollow CSEP-PATH protocols exactly; train and practise; use the same tester for retests
EquipmentUncalibrated sphygmomanometer, dynamometer or ergometer; wrong cuff size; worn measuring tapeCalibration and maintenance schedules; correct cuff sizing; check zero before use
EnvironmentHeat, humidity, noise, distractions, time of dayControlled room temperature; quiet space; test at a similar time of day
ProtocolChanging order of tests, rest periods, warm-up or footwear between visitsUse the same order, warm-up, clothing and footwear each time

Why it matters: measurement error can lead to wrong safety decisions, such as missing an elevated blood pressure. It can produce misleading Health Benefit Ratings, and it can show false progress or false decline that hurts client motivation and trust. Accurate measurement is also part of the standard of care a court would expect from a qualified exercise professional.

Loading diagram...
Reliability, Validity and Error
Test Your Knowledge

Two trainers independently time the same client's back extension test and record 92 seconds and 104 seconds. Which measurement property does this disagreement mainly reflect?

A

Criterion validity against a laboratory standard

B

Construct validity of the endurance test

C

Measurement sensitivity of the stopwatch

D

Objectivity (inter-rater reliability)

Test Your Knowledge

A bathroom-style scale consistently reads 2.0 kg higher than a calibrated medical scale every time the same client steps on it. How should this scale be described?

A

Valid but not reliable

B

Reliable but not valid

C

Neither reliable nor sensitive

D

Both valid and reliable

Test Your Knowledge

A client's sit-and-reach improves from 31.0 cm to 31.5 cm after four weeks. What is the most appropriate interpretation?

A

A meaningful improvement, because any recorded increase in reach must be a real change.

B

A decline, because genuine flexibility gains over four weeks are usually much larger.

C

A change too small to call real: it equals the 0.5 cm recording precision of the protocol.

D

Proof that the sit-and-reach test is invalid for this client and should be replaced.

Sections you finish are checked off in the contents.