1.5 Evaluating Work Capacity Across Broad Time and Modal Domains

Key Takeaways

  • CrossFit defines fitness as work capacity across broad time and modal domains, so evaluating an athlete means measuring power output at several durations and in several modalities — not one test.
  • Work is force times distance; power is work divided by time. A benchmark score converts to average power, which is what makes different workouts comparable.
  • A defensible assessment battery samples at least three time domains (roughly under 3 minutes, 5-15 minutes, and over 20 minutes) and all three metabolic pathways.
  • Gaps show up as a jagged profile: an athlete who is strong in short heavy efforts but collapses past 12 minutes has an aerobic deficiency, and the programming answer is more aerobic work, not more strength.
  • Retest conditions must match test conditions — same movements, same standards, same equipment, same scaling — or the comparison is meaningless.
Last updated: August 2026

Work Capacity Is the Measurement CrossFit Actually Cares About

CrossFit's definition of fitness is increased work capacity across broad time and modal domains. Every word of that definition is a measurement instruction. Work capacity means you must measure work, not effort. Broad time domains means one test is never enough. Broad modal domains means the tests must span weightlifting, gymnastics, and monostructural work. Domain 1 of the CCFT outline asks you to "evaluate athlete's work capacity" as a distinct task from assessing whether they can perform a functional movement — capability is can you do it, capacity is how much can you do.

The Physics You Are Expected to Use

  • Work = Force x Distance. Moving a 95 lb barbell overhead through 4 feet does 380 ft-lb of work per rep.
  • Power = Work / Time. Doing that same rep in 2 seconds produces 190 ft-lb/s; doing it in 4 seconds produces 95 ft-lb/s. Same work, half the power.
  • Intensity, in CrossFit's usage, is exactly power. This is why the methodology treats intensity as objectively measurable and independent of how hard something felt.

Worked example: an athlete completes 21-15-9 thrusters at 95 lb and pull-ups ("Fran") in 5:00. The barbell work alone is 45 reps x 95 lb x 4 ft = 17,100 ft-lb. Divided by 300 seconds, that is 57 ft-lb/s of average barbell power. The same athlete a year later finishes in 3:20 (200 s), producing 85.5 ft-lb/s — a 50% increase in average power at the same load. That number, not a subjective sense of fitness, is the evidence of adaptation.

Building a Defensible Assessment Battery

A single benchmark is a data point; a battery is an assessment. Sample the time domains and the modalities deliberately:

Time domainDominant pathwayRepresentative tests
0-30 secondsPhosphagen1RM back squat, 1RM clean & jerk, max unbroken strict pull-ups, 100 m sprint
30 s - 3 minGlycolytic500 m row, "Fran", 40 unbroken double-unders, 1 min max burpees
5-15 minMixed glycolytic/oxidative"Helen", "Cindy" (20 min variant), 2,000 m row, 1 mile run
20 min +Oxidative5 km run, 30 min AMRAP, "Murph", 10 km row

Across those, make sure you have at least one pure weightlifting test, one pure gymnastics test, one pure monostructural test, and two or three mixed-modal couplets or triplets. Six to eight benchmarks re-tested on a rolling basis is realistic for a group programme; twelve is realistic for individual design.

Reading the Profile

Plot the athlete's percentile or relative rank across the battery and look at the shape, not the average. Three common shapes and their programming answers:

  1. Strong-short, weak-long. Excellent 1RMs and sub-3-minute scores; times fall off a cliff past 12 minutes. This is an aerobic deficiency. The answer is more time at conversational pace and more 20+ minute work — not more strength cycles, which is what the athlete will want.
  2. Strong-long, weak-heavy. Good engine, poor absolute strength; the athlete is limited by load in every barbell workout. The answer is a dedicated strength block with reduced conditioning volume, accepting a temporary dip in the long-domain scores.
  3. Modally jagged. Fine on the rower and the barbell, falls apart on gymnastics. The answer is skill volume at low intensity — accumulated strict work, hollow/arch positions, and progressions — before adding gymnastics to metcons.

The CrossFit prescription is deliberately biased toward the worst domain: work your weaknesses. An athlete who only trains what they are good at gets a narrower, not broader, capacity.

Test Conditions and Valid Retesting

A retest is only evidence if the conditions match. Fix and record:

  • Movement standards — depth, lockout, chin over bar, hip crease below knee. If standards drifted, the score drifted.
  • Loads and implements — the same bar, the same rower with the same damper setting and drag factor, the same box height.
  • Scaling — a scaled retest compares only against the identical scaled version, and must be labelled as such.
  • Context — time of day, position in the training week, and what the athlete did in the preceding 48 hours. Retesting "Fran" the day after a heavy squat session is not a fair sample.
  • Interval — 8 to 12 weeks is the usual retest window. Retesting monthly produces noise; retesting yearly produces no feedback.

Document scores in a system the athlete can see. Visible, comparable data is both an assessment instrument and the single most reliable retention tool a coach has, because it makes invisible progress visible.

Sample Athlete Work-Capacity Profile (percentile rank vs gym population)
Test Your Knowledge

An athlete completes 45 reps of a 135 lb barbell moved 4 feet in 6 minutes. In a retest they complete the identical work in 4 minutes. What has changed?

A
B
C
D
Test Your Knowledge

An athlete ranks in the 85th percentile on 1RM lifts and sub-3-minute benchmarks but in the 20th percentile on every test longer than 20 minutes. What is the most appropriate programming response?

A
B
C
D
Test Your Knowledge

A coach retests 'Helen' and reports a 90-second improvement. Which single detail would most undermine the validity of that comparison?

A
B
C
D