5.4 Sampling Methods, Margin of Error, Bias, and Drawing Inferences

Key Takeaways

  • A parameter is a fixed numerical characteristic of an entire population, whereas a statistic is an empirical value calculated from a sample used to estimate the unknown population parameter.

  • Probability sampling designs (Simple Random, Stratified, Cluster, Systematic) utilize objective random mechanisms to eliminate selection bias; non-probability designs (Convenience, Voluntary Response) introduce systematic bias and cannot be generalized.

  • Observational studies can identify statistical correlations and associations, but only randomized controlled experiments with random assignment to treatments can establish cause-and-effect relationships.

  • Confounding variables are unmeasured extraneous factors correlated with both the explanatory and response variables that produce misleading or spurious associations in observational studies.

  • Margin of error reflects sampling variability due to chance; precision is primarily governed by sample size n (inversely proportional to √n), not the total population size N.

Last updated: September 2026

Sampling Methods, Margin of Error, Bias, and Drawing Inferences

OpenExamPrep provides this probabilistic and statistical reasoning review to help students master sampling techniques, study design, and statistical inference for the Texas Success Initiative Assessment 2.0 (TSIA2) Mathematics section. Interpreting statistical claims, distinguishing correlation from causation, and identifying bias are essential competencies for academic coursework and informed analysis of modern data.


1. Populations vs. Samples: Parameters vs. Statistics

Statistical inquiry seeks to understand characteristics of broad groups by examining representative subsets.

  • Population: The complete collection of all individuals, items, or measurements of interest (population size denoted N).
  • Sample: A representative subset of individuals selected from the population (sample size denoted n).
  • Parameter: A fixed numerical value that describes an entire population. Parameters are typically constant but unknown in practice because measuring every individual in a large population is impractical or impossible. Parameters are conventionally represented by Greek letters:
    • Population Mean: μ
    • Population Standard Deviation: σ
    • Population Proportion: p
  • Statistic: A numerical value computed directly from sample data. Statistics vary from sample to sample due to chance (sampling variability) and serve as point estimates of unknown population parameters. Statistics are represented by Roman letters:
    • Sample Mean: x̄
    • Sample Standard Deviation: s
    • Sample Proportion: p̂
Memory Anchor:
  Population ⟷ Parameter (both start with P)
  Sample     ⟷ Statistic (both start with S)

2. Sampling Methodologies

The method used to select a sample determines whether the sample statistics can be legitimately generalized to the target population.

Probability Sampling Methods (Valid for Generalization)

Probability sampling methods utilize objective chance mechanisms so that every individual has a known, non-zero probability of selection:

  1. Simple Random Sampling (SRS):
    • Every individual in the population has an equal chance of being selected, and every possible sample of size n has an equal probability of being chosen.
    • Implemented using random number generators or mechanical drawing from a thoroughly mixed pool.
  2. Stratified Random Sampling:
    • The population is divided into mutually exclusive, homogeneous subgroups called strata based on an important characteristic (e.g., academic major, age bracket, county).
    • A separate Simple Random Sample is drawn from every stratum proportionally.
    • Advantage: Guarantees representation of small minority subgroups and reduces sampling variability.
  3. Cluster Sampling:
    • The population is naturally divided into diverse, heterogeneous geographic or organizational units called clusters (e.g., city blocks, high schools, hospital wards).
    • A random selection of entire clusters is chosen; every individual within each selected cluster is surveyed.
    • Advantage: Highly practical and cost-effective when surveying geographically dispersed populations.
  4. Systematic Sampling:
    • Individuals are selected from an ordered list at regular, fixed intervals (every kth individual), beginning from a randomly chosen starting point between 1 and k.
    • Advantage: Easy to execute in continuous operational environments (e.g., inspecting every 20th item on an assembly line).

Non-Probability Sampling Methods (High Risk of Bias; Non-Generalizable)

  1. Convenience Sampling:
    • Selecting individuals who are easiest to contact or most readily accessible (e.g., interviewing students walking past a campus library at 10:00 AM).
  2. Voluntary Response Sampling:
    • A sample consisting entirely of self-selected volunteers who respond to an open public invitation (e.g., call-in radio polls, social media surveys).

Comprehensive Sampling Comparison Table

Sampling TechniqueSelection ProtocolPrimary AdvantageMajor Limitation / Vulnerability
Simple Random (SRS)Completely random lottery from entire sampling frameUnbiased; simple theoryRequires a complete, up-to-date population roster
Stratified RandomDivide into strata; SRS taken within every stratumGuarantees representation of all subgroupsRequires prior demographic knowledge of individuals
Cluster SamplingRandomly choose entire clusters; survey all membersLow administrative and travel costsHigh variance if clusters are internally homogeneous
Systematic SamplingSelect every kth individual from an ordered listOperational simplicity in flow processesSevere bias if list contains hidden periodic cycles
ConvenienceSelect whoever is readily accessibleInexpensive, rapid executionSevere selection bias; cannot generalize to population
Voluntary ResponseOpen call where participants choose to respondHigh participant engagementStrong bias toward individuals with extreme negative views

3. Sources of Statistical Bias

Bias is a systematic error in study design or execution that favors certain outcomes over others, causing sample statistics to consistently overestimate or underestimate the true population parameter.

  1. Selection Bias (Undercoverage): Occurs when the sampling frame systematically excludes or underrepresents certain segments of the target population. (e.g., An internet survey about municipal transit excludes low-income residents who lack broadband access).
  2. Nonresponse Bias: Occurs when a significant proportion of selected individuals fail or refuse to respond. If non-respondents differ systematically in their opinions or behaviors from those who respond, the resulting data are biased.
  3. Response Bias: Occurs when participants provide inaccurate, untruthful, or influenced answers:
    • Leading Question Wording: Questions phrased with loaded language that prompts a specific answer (e.g., "Do you agree that wasteful campus fees should be cut?").
    • Social Desirability Bias: Respondents hesitate to admit illegal, embarrassing, or socially disapproved behaviors to an interviewer.
  4. Voluntary Response Bias: Arises when sample members self-select; individuals with strong, passionate, or negative grievances are far more motivated to participate than the general public.

⚠️ Critical Rule: Increasing sample size does NOT cure bias! A biased voluntary response poll of 100,000 participants is vastly inferior to an unbiased Simple Random Sample of 400 individuals.


4. Study Design: Observational Studies vs. Randomized Controlled Experiments

A central distinction on the TSIA2 is understanding what type of conclusion is mathematically justified based on the study design:

                        Study Design Conclusions
                                   │
         ┌─────────────────────────┴─────────────────────────┐
         ▼                                                   ▼
Observational Study                                Randomized Experiment
- No treatments assigned                           - Treatments randomly assigned
- Observe existing traits                          - Control group / Placebo used
         │                                                   │
         ▼                                                   ▼
Establishes ASSOCIATION / CORRELATION              Establishes CAUSE-AND-EFFECT (Causation)

Observational Studies

In an observational study, researchers measure variables of interest without assigning treatments or attempting to influence responses.

  • Confounding Variables (Lurking Variables): An unmeasured extraneous factor that is correlated with both the explanatory variable and the response variable, creating a misleading or spurious association.
    • Classic Example: Ice cream sales and drowning incidents are strongly positively correlated. Eating ice cream does not cause drowning! The confounding variable is outdoor temperature (summer weather causes both an increase in ice cream consumption and an increase in swimming activity).
  • Limitation: Observational studies can establish correlation or association, but they can NEVER establish causation.

Randomized Controlled Experiments

In an experiment, researchers deliberately impose specific treatments on experimental units to observe their effects.

  • Principles of Experimental Design:
    1. Random Assignment (Randomization): Subjects are randomly assigned to treatment and control groups. Random assignment balances known and unknown confounding variables across groups, isolating the treatment as the sole cause of differences.
    2. Control Group & Placebo: A control group receiving an inactive treatment (placebo) establishes a baseline to measure the true physiological effect against psychological expectations (placebo effect).
    3. Replication: Administering treatments to a sufficiently large number of experimental units reduces the impact of chance variation.
    4. Blinding:
      • Single-Blind: Subjects do not know which treatment they are receiving.
      • Double-Blind: Neither the subjects nor the evaluating researchers know which subjects received which treatment, eliminating experimenter bias.

5. Statistical Inference, Margin of Error, and Generalization

Statistical inference involves using sample data to make generalized statements about unknown population parameters.

Margin of Error (MOE) and Confidence Intervals

Because random samples fluctuate due to chance, point estimates are reported with a margin of error:

Confidence Interval = Point Estimate ± Margin of Error
  • If a poll reports that 54% of college students favor an initiative with a margin of error of ±3% at a 95% confidence level:
    • Interval: 54% ± 3% = [51%, 57%]
    • Interpretation: We are 95% confident that the true population proportion of all college students favoring the initiative lies between 51% and 57%.

The Inverse Square-Root Relationship of Sample Size

For estimating a population proportion, the margin of error is inversely proportional to the square root of the sample size n:

Margin of Error (MOE) ≈ 1 ÷ √n
  • To cut the margin of error in half (e.g., from ±4% down to ±2%), the sample size n must be quadrupled (multiplied by 4): √4n = 2√n in the denominator cuts the margin of error in half.
  • Population Size Independence: As long as the population is at least 10 to 20 times larger than the sample, the total population size N has virtually no impact on the margin of error. A random sample of 1,000 voters yields approximately the same margin of error whether drawn from Houston, Texas or the entire United States.

Valid Generalization Principles

  1. Results can be generalized only to the population from which the random sample was selected.
  2. If participants were not randomly selected (convenience or voluntary response), conclusions apply strictly to the participants themselves.
  3. Cause-and-effect claims require random assignment to treatments, not merely random sampling from a population.

6. TSIA2 Exam Traps & Strategic Checkpoints

  • Trap 1: Correlation vs. Causation: In any question describing a survey or observational study, eliminate answer choices claiming that one factor "causes", "prevents", or "cures" another.
  • Trap 2: Stratified vs. Cluster Sampling Distinction:
    • Stratified Sampling: Sample some individuals from all groups (strata).
    • Cluster Sampling: Sample all individuals from some groups (clusters).
  • Trap 3: The Large Sample Illusion: Do not select voluntary internet polls as reliable simply because they have millions of respondents. Bias cannot be overcome by sample size.
  • Trap 4: Margin of Error and Population Size: Precision depends on sample size n, not population size N. A poll of 1,000 Texans has the same margin of error as a poll of 1,000 Rhode Islanders.
Loading diagram...
Study Design Framework: Observational vs. Experimental Conclusions
Margin of Error (%) vs. Sample Size (n): Inverse Square-Root Relationship
Test Your Knowledge

A medical research team conducts an observational study tracking 5,000 adults over ten years and discovers that individuals who drink at least three cups of green tea daily have a 25% lower incidence of cardiovascular disease than non-tea drinkers. Which conclusion is statistically valid?

A

Drinking three cups of green tea daily causes a direct reduction in cardiovascular disease risk

B

The study proves that non-tea drinkers can prevent heart disease simply by adding green tea to their diet

C

The sample size is too small to draw any conclusions regarding green tea and cardiovascular health

D

There is an association between green tea consumption and lower cardiovascular disease incidence, but confounding variables prevent establishing a causal relationship

Test Your Knowledge

A university administrator wants to estimate student satisfaction with campus dining services. The student body consists of 6,000 freshmen, 5,000 sophomores, 4,500 juniors, and 4,500 seniors. The administrator randomly selects 60 freshmen, 50 sophomores, 45 juniors, and 45 seniors to survey. Which sampling method was utilized?

A

Stratified random sampling

B

Cluster sampling

C

Systematic sampling

D

Convenience sampling

Test Your Knowledge

A pre-election telephone survey of 1,024 randomly selected registered voters indicates that 53% support a proposed municipal bond measure, with an announced margin of error of ±3.1 percentage points at a 95% confidence level. How should this result be interpreted?

A

The bond measure is guaranteed to pass because 53% is strictly greater than 50%

B

We are 95% confident that the true proportion of all registered voters supporting the bond measure lies between 49.9% and 56.1%

C

Exactly 3.1% of the voters surveyed gave dishonest or incorrect responses

D

Repeating the identical survey in a larger city would require ten times as many respondents to achieve the same margin of error

Test Your Knowledge

A research pollster wishes to reduce the margin of error of a nationwide public opinion survey from ±4 percentage points down to ±2 percentage points at the same confidence level. How must the sample size n be altered?

A

The sample size must be doubled (multiplied by 2)

B

The sample size must be cut in half (divided by 2)

C

The sample size must be quadrupled (multiplied by 4)

D

The sample size must be increased by a factor of 16

Sections you finish are checked off in the contents.