8.5 Sampling, MSA & Reliability Terminology
Key Takeaways
- Measurement planning begins with the decision, requirement, operational definition, and risk—not with an available instrument.
- Random sampling supports unbiased inference; stratified sampling deliberately represents important subgroups.
- Accuracy concerns closeness to a reference value, while precision concerns consistency of repeated readings.
- Gauge R&R separates repeatability from reproducibility to evaluate measurement-system variation.
- CMQ/OE reliability questions require terminology and concepts, not reliability calculations.
Measure for a decision, not because data are available
Measurement is useful only when it supports a decision. IV.C asks managers to understand when, what, and how to measure. Start with the business or quality question: Are customers receiving conforming service? Is a process meeting a cycle-time commitment? Is a corrective action reducing recurrence? Then define the characteristic, the population, the unit, timing, method, owner, and response when the result is outside expectations.
When to measure depends on risk and process knowledge. Measure before a high-risk decision, at a control point where prevention is possible, when performance can change, and often enough to detect a meaningful shift before customers are affected. Overmeasurement can create delay and cost; undermeasurement can allow undetected defects or deteriorating service. A manager balances the cost of information against the cost of a wrong decision.
What to measure flows from requirements and process goals. Customer-facing measures may include complete-and-accurate orders, response time, defect-free delivery, complaint resolution, or ease of use. Process measures may include first-pass yield, queue time, rework, setup time, downtime, and error rate. Use a balanced set: a speed measure alone can encourage rushed, inaccurate work; an output count alone can encourage overproduction.
How to measure means creating an operational definition and a capable method. An operational definition states precisely what is included, excluded, counted, classified, and recorded. For a "late order" metric, define the promise date, time zone, shipment event, treatment of customer holds, and data source. Different people applying the same definition should reach the same conclusion. The method must be practical, safe, timely, calibrated where applicable, and reliable enough for the decision.
| Planning question | Example of a complete answer | Why it matters |
|---|---|---|
| Decision | Should the revised intake process be standardized? | Prevents collecting data without a use. |
| Characteristic | Percentage of requests completed correctly on the first pass | Connects the measure to quality. |
| Population and unit | All requests closed this month; one request is one unit | Defines what claims the result can support. |
| Operational definition | Correct means all required fields, approvals, and customer commitments are satisfied | Improves consistency of classification. |
| Collection plan | Weekly random sample, independent reviewer, common form | Makes the data reproducible and auditable. |
| Response | Investigate a sustained decline or a point beyond the control limits | Connects monitoring to action. |
Sampling: represent the population fairly
Sampling is used when measuring every item is impractical, destructive, too slow, or unnecessary for the decision. A sample supports inference about a defined population only when the selection method fits that purpose. The central risk is bias: a convenient sample may systematically exclude the conditions where performance is poor.
A random sample gives each relevant population unit a known, nonzero chance of selection, commonly an equal chance in simple random sampling. Random selection helps prevent conscious or unconscious cherry-picking. It does not mean grabbing whatever happens to be nearby, selecting the first ten records, or asking volunteers. A random-number generator applied to a complete sampling frame is one practical approach.
A stratified sample divides the population into meaningful, relatively homogeneous subgroups—called strata—and samples from each. Strata could be shift, region, product family, customer segment, supplier, channel, or risk class. Stratification is appropriate when an overall random sample might underrepresent a small but important subgroup or when managers need subgroup estimates. For example, a national service organization should not sample only its largest region if it must understand performance in every region. The selection within each stratum should still be random.
| Sampling approach | Best use | Caution |
|---|---|---|
| Simple random | Population is reasonably uniform and one overall estimate is needed | Requires a usable complete sampling frame. |
| Stratified random | Important segments, shifts, products, or locations must be represented | Combine results using the correct population weights when estimating an overall rate. |
| Convenience | Rapid exploratory learning only | Cannot support confident population conclusions because selection bias is likely. |
| Census / 100% inspection | Risk is very high, population is small, or automation makes it feasible | Inspecting everything still does not overcome a poor measurement method. |
A CMQ/OE manager should also consider sample size, timing, and independence. A sample collected only on day shift or only after an improvement announcement can be biased even if records were randomly selected within that narrow window. Sampling plans should state the population, frame, selection method, sample size rationale, frequency, stratification, and how results will be used. If data are used to judge teams, independent collection or review can reduce conflict of interest.
Measurement system analysis (MSA)
A measurement system includes the instrument, software, procedure, fixture, environment, appraiser, training, sampling method, and data-recording process. Measurement system analysis (MSA) asks whether that system produces data good enough for the intended use. A sophisticated instrument does not guarantee good data if operators interpret the method differently, the fixture is unstable, the system is poorly calibrated, or the data-entry rule is unclear.
The most frequently tested terms are related but not interchangeable:
| Term | Meaning | Example |
|---|---|---|
| Accuracy | Closeness of a measured value to an accepted reference or true value | A scale reading 100.0 g for a certified 100.0 g weight is accurate. |
| Precision | Closeness of repeated measurements to one another | A scale repeatedly reading 98.0 g may be precise but inaccurate. |
| Bias | Systematic difference between observed average and a reference value | A thermometer consistently reads 1.5°C too high. |
| Linearity | Change in bias across the measurement range | A device is accurate at low values but increasingly high at large values. |
| Repeatability | Variation when the same appraiser measures the same item repeatedly using the same method and equipment | One inspector repeatedly measures one part differently. |
| Reproducibility | Variation in average results among appraisers using the same method and equipment | Three inspectors obtain materially different values on the same parts. |
Accuracy and precision produce four possible conditions. A system can be both accurate and precise, precise but biased, accurate on average but highly variable, or neither. For management decisions about product acceptance or improvement, the desired condition is both accurate and precise. A calibration check may reveal bias against a reference; repeated trials may reveal poor precision. Do not use the terms as synonyms.
Bias is a systematic offset. If a scale consistently reads low, adjusting or calibrating it may address the issue. Linearity asks whether that bias is consistent throughout the operating range. A device can appear satisfactory when checked only at one reference point and still be unacceptable at the high end of its use range. Therefore, assess the measurement range relevant to the process decision.
A gauge repeatability and reproducibility study, usually called gauge R&R, evaluates two major components of measurement variation. Repeatability is equipment variation under the same operator and conditions. Reproducibility is appraiser-to-appraiser variation. A typical variable-data study has multiple appraisers measure multiple parts more than once in randomized order. The study design helps separate part-to-part differences from the measurement-system differences. For attribute data, an agreement study may assess whether appraisers consistently classify items as acceptable or defective.
Gauge R&R is not merely a statistic to report. A high measurement-system contribution means apparent process variation or apparent improvement may actually be measurement noise. Management responses can include clarifying the operational definition, improving fixtures, controlling environment, calibrating equipment, training appraisers, simplifying a subjective rating scale, or selecting a more capable method. Improve the measurement system before using its output to punish a team or make a costly process decision.
Reliability terminology: definitions, not calculations
Reliability is the probability that an item, system, or process performs its intended function without failure for a specified time under stated conditions. Reliability is contextual: a device reliable in a clean indoor environment may not be reliable in heat, vibration, or moisture. A reliable process consistently delivers its intended output over time under expected operating conditions. It is not simply a process that produced one good batch.
The bathtub curve describes a common pattern of failure rate over a product life cycle:
- Infant mortality: early failures are relatively high because of latent manufacturing defects, installation errors, weak components, or early-use issues. Screening, burn-in where appropriate, process control, and design improvement may reduce them.
- Useful life: failures occur at a relatively low, approximately constant rate. Random stresses and chance events dominate.
- Wear-out: failure rate rises as components age, fatigue, corrode, wear, or degrade. Preventive replacement and life-cycle planning become important.
Mean time between failures (MTBF) is the average operating time between failures for repairable items. It is commonly used as an indicator of reliability for a repairable system. Mean time to repair (MTTR) is the average time required to restore a repairable item after failure, often including diagnosis, repair, testing, and return to service according to the organization’s definition. MTTR is primarily a maintainability or restore-time concept, though it affects availability and customer experience. Define start and end points consistently before comparing MTTR.
Do not infer that a high MTBF automatically means low downtime: a system can fail infrequently but take a long time to repair. Likewise, a low MTTR does not mean failures are rare. A quality manager chooses separate measures for failure frequency, repair time, availability, customer disruption, and cost when those distinctions matter.
Exam boundary: the CMQ/OE Body of Knowledge tests reliability terminology and concepts in this area. You should recognize infant mortality, the bathtub curve, MTBF, MTTR, and process reliability, but reliability calculations are not tested. Focus your study time on defining the terms, selecting the appropriate measure in a scenario, and explaining what a failure pattern implies for quality management.
Scenario check
A hospital reports that a new diagnostic device has many failures during its first month, then stabilizes. The correct concept is infant mortality, not wear-out. The manager should investigate installation, early component defects, training, and incoming quality. If the same device fails more frequently after several years of heavy use, wear-out is the more plausible pattern. Measurement and reliability data should be defined and sampled in a way that lets leaders distinguish these causes rather than reacting to anecdote.
A company needs to compare on-time performance for each of four regions, including a small region that represents only 4% of orders. Which sampling plan is most appropriate?
Three trained inspectors classify the same borderline samples differently even though each inspector is consistent when repeating their own assessments. Which MSA issue is most directly indicated?
What does MTTR describe for a repairable system?