Sampling Methodologies: Random, Stratified, Convenience, and Targeted

Key Takeaways

  • Probability sampling methods (simple random and stratified random) ensure every item in the population universe has a known, non-zero chance of selection, making them legally and statistically defensible for financial extrapolation.
  • Stratified random sampling divides a heterogenous population universe into non-overlapping sub-groups (strata) based on dollar values or code categories, significantly reducing variance and standard error.
  • Targeted (focused) sampling selects claims based on high-risk parameters, making it ideal for root-cause discovery and provider education, but strictly prohibiting financial extrapolation across the general universe.
  • Convenience sampling relies on accessible records, introduces extreme selection bias, and is invalid for formal compliance reporting or payer disputes.
  • HHS-OIG RAT-STATS software is the official statistical tool required for generating random samples and determining confidence intervals in federal self-disclosure protocols and corporate integrity agreements.
Last updated: July 2026

Sampling Methodologies: Random, Stratified, Convenience, and Targeted

Core Principle: The choice of sampling methodology determines whether audit findings can be extrapolated across an entire population frame. Probability sampling enables valid financial projections, whereas non-probability sampling restricts findings strictly to the audited charts.

In medical record auditing, reviewing 100% of claims submitted by a practice or facility over a multi-year period is rarely practical due to administrative costs and time constraints. Auditors must select a subset of claims—a sample—to evaluate documentation integrity. However, the mathematical method used to draw that sample dictates the legal and statistical validity of the audit's conclusions. The Certified Professional Medical Auditor must master the structural differences, selection mechanics, advantages, and regulatory limitations of probability and non-probability sampling methodologies.


Probability vs. Non-Probability Sampling

All audit sampling techniques fall into two primary classifications: Probability (Statistical) Sampling and Non-Probability (Non-Statistical / Judgmental) Sampling.

1. Probability (Statistical) Sampling

In probability sampling, every sampling unit (e.g., claim, claim line, or patient record) in the defined population universe ($N$) has a known, non-zero probability of being selected. Selection is executed using objective randomization tools without human bias.

  • Key Characteristic: Enables mathematical measurement of sampling error, standard error, and confidence intervals.
  • Financial Extrapolation: Permitted. Findings from a statistically valid random sample (SVRS) can be projected across the entire population universe to calculate total estimated overpayments.

2. Non-Probability (Non-Statistical) Sampling

In non-probability sampling, items are selected based on auditor judgment, risk criteria, or physical convenience. The probability of any specific claim being chosen is unknown.

  • Key Characteristic: Subject to selection bias; cannot measure standard error or statistical confidence.
  • Financial Extrapolation: Prohibited. Overpayment calculations are legally restricted only to the specific claims audited. Extrapolating non-probability sample error rates onto an entire population frame is statistically invalid and legally defensible by the provider.
FeatureProbability SamplingNon-Probability Sampling
Selection CriteriaRandomization (equal/known chance)Auditor judgment, risk rules, or convenience
Selection BiasEliminatedPresent / Inherent
Standard Error CalculationMathematically measurableCannot be calculated
Extrapolation Permitted?YES (with valid confidence limits)NO (restricted to audited sample dollars)
Primary Use CaseGovernment audits, CIAs, overpayment quantificationProvider education, probe reviews, root-cause discovery

Probability Sampling Techniques: Simple Random and Stratified Random

1. Simple Random Sampling (SRS)

Simple Random Sampling is the foundational statistical technique where every claim in the population universe has an identical probability of selection.

  • Selection Mechanics: The auditor assigns a sequential numerical identifier to every claim in the population frame ($1$ to $N$). A random number generator (such as HHS-OIG's RAT-STATS software) selects $n$ unique numbers.
  • Systematic Sampling Alternative: A variant of SRS where the auditor calculates a sampling interval $k = N / n$. Starting at a randomly selected claim between $1$ and $k$, the auditor selects every $k$-th claim throughout the list.
  • Advantages: Completely unbiased, straightforward to execute, and highly defensible in administrative appeals.
  • Limitations: In a heterogenous population frame containing wide financial variances (e.g., mixing $50 routine office visits with $30,000 inpatient surgical procedures), simple random sampling can result in a high standard error unless the sample size is very large.

2. Stratified Random Sampling

Stratified Random Sampling addresses population heterogeneity by dividing the population universe frame ($N$) into non-overlapping, mutually exclusive sub-groups called strata based on specific attributes. The auditor then conducts independent simple random sampling within each stratum.

  • Common Stratification Criteria:
    • Dollar Ranges: Stratum 1: claims <$500; Stratum 2: claims $500–$4,999; Stratum 3: claims >= $5,000.
    • Service Categories: Stratum A: Evaluation & Management; Stratum B: Surgical Procedures; Stratum C: Diagnostic Imaging.
    • Provider Specialty or Location: Stratified by individual NPIs within a large group practice.
[ Total Population Universe: 5,000 Claims (Heterogenous) ]
                           |
      +--------------------+--------------------+
      |                                         |
[ Stratum 1: E/M Codes ]               [ Stratum 2: Major Surgeries ]
  (N1 = 4,000 claims)                     (N2 = 1,000 claims)
      |                                         |
[ Simple Random Sample n1=80 ]         [ Simple Random Sample n2=40 ]
  • Advantages:
    • Reduces Standard Error & Variance: Grouping similar claim values together minimizes sampling variance within each stratum, producing a tighter confidence interval.
    • Ensures High-Risk Representation: Guarantees that rare, high-dollar surgical claims are adequately sampled rather than being missed by a simple random pull.
    • Allows Sub-Group Analysis: Enables the auditor to report distinct error rates for different procedure types.

Non-Probability Sampling Techniques: Targeted and Convenience

1. Targeted / Focused / Purposive Sampling

In targeted sampling, the auditor purposefully selects claims that meet specific high-risk criteria identified through data mining, billing software rules, or external compliance alerts.

  • Selection Criteria: Pulling all claims appended with Modifier 25, all claims billed with CPT 99215 on a single date of service, or claims submitted by a provider under a corrective action plan.
  • Primary Purpose: Excellent for preliminary probe audits, discovering root causes of coding errors, and conducting focused provider training.
  • Critical Limitation: Targeted sampling intentionally selects known risk items, creating severe selection bias. It is legally impermissible to extrapolate the error rate from a targeted sample across the broader general billing universe, as doing so would project high-risk error frequencies onto un-audited, low-risk services.

2. Convenience / Judgmental Sampling

In convenience sampling, claims are selected purely based on ease of access or auditor availability.

  • Selection Examples: Auditing the last 15 paper charts sitting on the clinic desk, or selecting charts for patients seen on a single Monday morning because the charts were physically available.
  • Vulnerabilities: Suffers from extreme selection bias. Charts available on a Monday morning may represent specific walk-in patients or a single attending physician, failing to represent the true diversity of the practice.
  • Regulatory Status: Strictly prohibited in formal compliance audits, payer dispute proceedings, and OIG self-disclosure reporting.

Method Selection Matrix & OIG RAT-STATS Guidelines

When selecting a sampling methodology, CPMA auditors evaluate objective requirements against regulatory standards:

Sampling MethodRandom Selection?Extrapolation Valid?Primary AdvantageMain Limitation
Simple RandomYesYesUnbiased; easy to explainHigh variance if claim dollars vary wildly
Stratified RandomYesYesMinimizes standard error & sample sizeRequires pre-audit claim data sorting
Targeted / FocusedNoNo (Sample dollars only)Efficient for root-cause error discoverySelection bias invalidates extrapolation
ConvenienceNoNoRequires zero preparationHighly biased; zero statistical credibility

HHS-OIG RAT-STATS Software

RAT-STATS is the free statistical software package provided by the HHS-OIG Office of Audit Services. It is the mandatory gold standard for generating random samples and evaluating audit results under OIG Corporate Integrity Agreements (CIAs) and the OIG Self-Disclosure Protocol (SDP).

  • Random Number Generator Module: Generates verifiable, reproducible random sample pulls using fixed seed numbers.
  • Sample Size Determination Module: Calculates exact sample size requirements based on target confidence levels (typically 90% or 95%), desired precision margins (e.g., +/- 5% or +/- 10%), and anticipated population standard deviation.
  • Unstratified & Stratified Framing: Supports both single-stage simple random sampling and complex multi-strata designs.
Test Your Knowledge

An auditor is planning a retrospective compliance audit for an ambulatory surgical center (ASC) covering 5,000 claims billed over the past calendar year. The claim population consists of 4,000 low-cost pain management injections (averaging $150 per claim) and 1,000 complex orthopedic surgical procedures (averaging $8,500 per claim). The goal of the audit is to calculate a statistically valid, highly precise financial overpayment estimate. Which sampling methodology should the auditor implement?

A
B
C
D
Test Your Knowledge

A clinic auditor conducts a focused audit of 40 claims where Modifier 25 was appended to an E/M code on the same day as a minor procedure. The 40 charts were specifically selected because data mining identified them as potential unbundling risks. The audit reveals an error rate of 25% in the sample, totaling $2,000 in improper payments. The auditor then multiplies this 25% error rate by the clinic's total annual billing of $1,000,000 across all services to demand a $250,000 refund. Why is this extrapolation method legally and statistically invalid?

A
B
C
D
Test Your Knowledge

A healthcare organization is preparing a submission under the HHS-OIG Self-Disclosure Protocol (SDP) after identifying systemic billing errors in its outpatient clinic. To comply with OIG requirements for quantifying total financial overpayment across a population frame of 12,000 claims, which tool and sampling procedure must the organization utilize?

A
B
C
D
Test Your Knowledge

During an internal review of a neurology practice, an auditor audits 15 patient charts pulled directly from the physical chart stack waiting on the front desk counter on a Tuesday morning. The auditor uses the findings from these 15 charts to issue a formal compliance rating for the entire 10-physician practice. Which type of sampling was performed, and what is its primary weakness?

A
B
C
D