6.5 Data Collection Plans and Common Pitfalls

Key Takeaways

  • A data collection plan specifies what, why, how, who, when, where, how many, and how the data will be recorded and stratified.
  • The three pitfalls ASQ names are undefined metrics, collectors untrained in the tools or unaware of how the data will be used, and unchecked seasonality effects.
  • Stratification factors must be captured at the moment of collection, because they cannot be reconstructed afterwards.
  • Check sheets are the standard structured collection instrument and should be designed so that recording is faster than not recording.
  • Data normalization converts raw counts into comparable rates, without which comparisons across units of different size are invalid.
Last updated: August 2026

The plan

A data collection plan is a table, one row per measure, answering eight questions.

FieldContentExample
WhatThe measure and its operational definitionFill weight in grams, measured at the check-weigher, at the moment of seal
WhyWhich project question it answersBaseline capability against 500 g +/- 8 g
HowInstrument, method, resolutionCheck-weigher #3, 0.1 g resolution, MSA completed 12 Aug
WhoNamed role, trained and certifiedLine QA technician, certified on the work instruction
WhenFrequency and timing5 consecutive units every 30 minutes
WhereLocation and stratificationLines 1-3, all shifts, both resin lots
How manySample size with its justification25 subgroups of 5 = 125 units, for control chart establishment
RecordingForm, field names, destination, retentionCheck sheet CS-114, keyed to the MES within one shift

Two fields are habitually skipped and both are costly. Why prevents collection of data nobody uses. Where, in the sense of stratification factors, cannot be added retrospectively.

Data integrity and accuracy

Distinguish three things that are frequently conflated:

  • Accuracy -- how close the measurement is to the true value. Addressed by calibration and by bias studies in MSA.
  • Precision -- how repeatable the measurement is. Addressed by gage R&R.
  • Integrity -- whether the record faithfully represents what happened. Addressed by collection design, training, and traceability.

Integrity failures are the ones a statistical analysis cannot detect. Transcription errors, entries made at the end of shift from memory, defaults left unchanged, unit mix-ups, and duplicated records all produce data that looks well-behaved and is wrong.

Practical safeguards:

  • Record at the point and moment of occurrence, never retrospectively.
  • Use fixed value lists rather than free text for categorical fields.
  • Include a unique identifier per record so duplicates are detectable.
  • Include the collector, the instrument ID, and the timestamp on every record.
  • Range-check and plausibility-check on entry; a fill weight of 5,000 g should be rejected at the keyboard.
  • Reconcile counts: units measured should tie to units produced.

The three pitfalls named by ASQ

1. Metrics not defined

If the metric is not operationally defined, different collectors record different things and the variation you measure is partly theirs. Two collectors working independently from the definition alone must classify the same items identically -- and attribute agreement analysis is the formal test of whether they do. Where a definition proves ambiguous, fix the definition and re-collect rather than adjudicating the disputed records.

2. Collectors untrained and unaware of how data will be used

The Body of Knowledge is specific that collectors must be trained in the tools and understand how the data will be used. The second half is the part usually skipped, and it matters for two reasons. People who know the purpose catch anomalies -- "this reading looks wrong, the gauge was just knocked" -- that a mechanical recorder would enter unchallenged. And people who do not know the purpose invent one, which is where fear-driven distortion begins. State explicitly that the study examines the process and not the individual.

3. Seasonality not checked

Seasonality here means any systematic time-related pattern: hour of day, day of week, shift, week of month, month of year, or campaign and lot cycles.

Two distinct failures follow from ignoring it. A baseline collected during an unrepresentative window gives a wrong baseline -- a call centre measured in the first week of January is not the same process as in the third week of June. And a before-and-after comparison that spans a seasonal change credits the season to the project. The remedies are to collect across at least one full cycle of the dominant period, to stratify by time factors, and where a full cycle is impossible, to compare like periods year over year rather than consecutive periods.

Stratification factors

Capture them at collection time. The standard candidates:

FactorWhy it matters
Machine, line, or stationEquipment differences are a leading variation source
Operator or teamMethod differences between people
ShiftStaffing, supervision, and support differ by shift
Time of day / day of weekWarm-up, fatigue, demand pattern
Material lot or supplierInput variation
Product, SKU, or customer segmentDifferent requirements and routings
Location or siteEnvironment, local practice

The rule: if you might later want to compare two groups, you must record which group each observation belongs to at the time you record it. Reconstructing "which shift produced this unit" from timestamps three months later is unreliable and sometimes impossible.

Check sheets

A check sheet is a pre-formatted form that structures recording at the point of collection. Common forms: the defect-type tally, the defect-location diagram (marks placed on a picture of the product), the defect cause sheet crossing cause categories with time, and the tally distribution sheet that builds a histogram as data arrives.

Design rules:

  • Make recording faster than not recording. A tick beats a written number; a written number beats a sentence.
  • Pre-print all categories, including "other" with a free-text prompt so novel modes are captured.
  • Include the stratification fields as pre-printed columns rather than relying on memory.
  • Pilot the sheet for one shift before full deployment. Every check sheet has an error in its first version.

Data normalization

Raw counts are not comparable across units of different size, volume, or exposure. Normalization converts them into rates:

Defect rate=DefectsOpportunities,DPU=DefectsUnits,Rate=EventsExposure\text{Defect rate} = \frac{\text{Defects}}{\text{Opportunities}}, \qquad DPU = \frac{\text{Defects}}{\text{Units}}, \qquad \text{Rate} = \frac{\text{Events}}{\text{Exposure}}

Plant A reporting 340 defects and plant B reporting 95 says nothing until you know plant A produced 120,000 units and plant B produced 11,000: the rates are 2.8 and 8.6 per thousand, and the smaller plant is three times worse. Choose the denominator that represents genuine exposure -- units, hours, opportunities, patient-days, transactions -- and use the same denominator on both sides of every comparison.

Before you start collecting

A short pre-flight checklist:

  1. Is the measurement system validated (MSA complete and passed)?
  2. Is every metric operationally defined and tested for agreement between collectors?
  3. Are the collectors trained, and do they know how the data will be used?
  4. Are all stratification factors on the form?
  5. Does the collection window cover at least one full cycle of the dominant time pattern?
  6. Is the sample size justified by a calculation rather than by convenience?
  7. Has the form been piloted for one shift?
  8. Is there a plan for missing and out-of-range values?
Test Your Knowledge

A team collects three weeks of baseline data on a retail returns process in mid-January and sets the project baseline from it. Which pitfall named in the Body of Knowledge has been risked?

A
B
C
D
Test Your Knowledge

Why must stratification factors such as shift, machine, and material lot be recorded at the moment of data collection?

A
B
C
D
Test Your Knowledge

Plant A reports 340 defects from 120,000 units and Plant B reports 95 defects from 11,000 units. What does normalization reveal?

A
B
C
D