15.1 Statistical Study Design, Random Sampling Techniques & Sources of Bias

Key Takeaways

  • Statistical investigations are partitioned into sample surveys, observational studies, and controlled experiments; only well-designed randomized comparative experiments can establish cause-and-effect relationships because randomization balances confounding variables across treatment groups.
  • Probability sampling designs—simple random, stratified, cluster, and systematic sampling—rely on objective chance mechanisms, whereas non-probability designs like voluntary response and convenience sampling suffer from fatal selection bias.
  • Stratified sampling samples from all homogeneous strata to minimize sampling variance, while cluster sampling surveys all individuals within a subset of randomly selected heterogeneous clusters for operational efficiency.
  • Systematic errors in survey research stem from selection bias and undercoverage, non-response bias among unreachable subjects, response bias from inaccurate reporting, and leading question wording.
  • The four foundational pillars of experimental design are comparison, randomization, control, and replication, supplemented by blinding protocols and randomized block designs to eliminate placebo effects and isolate treatment efficacy.
Last updated: September 2026

15.1 Statistical Study Design, Random Sampling Techniques & Sources of Bias

Taxonomy of Empirical Studies: Surveys, Observational Studies, and Experiments

Statistical inquiry begins with data collection. The methodology used to collect data dictates what conclusions can legitimately be drawn. The discipline classifies empirical investigations into three distinct paradigms: sample surveys, observational studies, and controlled experiments.

A sample survey collects data from a representative subset of a population to estimate unknown population parameters through structured questionnaires, interviews, or standardized physical measurements. In an ideal survey, the investigator records responses without attempting to influence subjects or alter their environment.

An observational study observes and measures variables of interest across individuals in their natural state without imposing any treatment or intervening in any way. Observational studies can be retrospective, where researchers examine historical records or past behaviors (such as reviewing medical archives to link dietary habits with cardiac health), or prospective, where cohorts are tracked forward in time to record emergent outcomes. Because the investigator does not assign subjects to conditions, observational studies are inherently vulnerable to confounding variables—extraneous factors that correlate simultaneously with both the explanatory variable and the response variable, obscuring whether the explanatory variable truly produces changes in the response.

A controlled experiment, by contrast, involves the deliberate imposition of one or more experimental treatments on subjects or experimental units to measure their subsequent responses. The defining criterion of an experiment is active intervention: researchers manipulate the explanatory variable (factor) to assess its direct impact on the response variable.

A fundamental tenet of statistics tested on the FTCE Mathematics 6-12 exam is that correlation does not imply causation. An observational study or sample survey can reveal strong associations between two variables, but it cannot establish cause-and-effect. For instance, an observational study might reveal a strong positive correlation between ice cream sales and drowning incidents; however, warmer summer temperatures act as a confounding variable driving increases in both. Only a well-designed, randomized comparative experiment can establish a causal relationship, because the random allocation of subjects to treatments balances out lurking and confounding variables across groups, isolating the treatment as the sole systematic difference.


Probability vs. Non-Probability Sampling Designs

To generalize conclusions from a sample to an entire target population, researchers must utilize probability sampling designs, in which every member of the population has a known, non-zero probability of selection determined by an objective chance mechanism.

  1. Simple Random Sample (SRS): An SRS of size $n$ is chosen in such a manner that every individual in the population has an equal probability of selection, and every possible subset of $n$ individuals has an equal chance of constituting the chosen sample. Common implementations include assigning integers to every member in a sampling frame and generating values with a pseudo-random number generator. While an SRS provides an unbiased benchmark, it can be logistically challenging or costly across geographically dispersed populations.

  2. Stratified Random Sample: The population is partitioned into mutually exclusive, internally homogeneous subgroups called strata based on a shared characteristic known or suspected to correlate with the variable of interest (e.g., grouping high school students by grade level: 9th, 10th, 11th, and 12th). An independent SRS is drawn from every stratum, and the sampled individuals are pooled. Stratification guarantees representation of key demographic subgroups, substantially reducing sampling variability across repeated samples.

  3. Cluster Sample: The population is divided into pre-existing, naturally occurring, internally heterogeneous groups called clusters (e.g., city blocks, school homerooms, or voting precincts). A random sample of clusters is selected, and all individuals within the chosen clusters are surveyed. Unlike stratified sampling—which samples some individuals from all groups—cluster sampling measures all individuals from some groups. Cluster sampling provides operational and travel efficiencies, though it yields higher sampling variability if clusters are internally homogeneous rather than representative mini-populations.

  4. Systematic Sample: From an ordered population list of size $N$, researchers select every $k$-th individual (where $k \approx N/n$) after randomly choosing a starting point between 1 and $k$. Systematic sampling is simple to administer in industrial quality control or customer exit surveys, but it introduces severe bias if the population list exhibits a periodic or cyclical pattern that coincides with $k$.

In sharp contrast, non-probability sampling designs rely on convenience or self-selection rather than chance mechanisms, introducing fatal systemic bias:

  • Convenience Sample: Researchers select individuals who are easiest to contact or recruit (e.g., questioning shoppers at a single shopping mall entrance on a weekday morning). Convenience samples systematically exclude large segments of the population.
  • Voluntary Response Sample: The sample consists of individuals who choose themselves by responding to an open invitation (e.g., online polls, call-in radio shows). Voluntary response samples overwhelmingly overrepresent individuals with intense, strongly held, or negative opinions.

Sources of Systematic Bias in Sampling

Bias refers to any systematic favoritism in data collection that causes sample statistics to consistently overestimate or underestimate the true population parameter.

  • Undercoverage (Selection Bias): Occurs when the sampling frame (the list of individuals from which the sample is drawn) fails to include certain population segments. For example, a telephone survey relying solely on registered landlines systematically undercovers younger and lower-income demographics who rely exclusively on mobile phones.
  • Non-Response Bias: Arises when selected individuals refuse to participate or cannot be contacted. If individuals who do not respond possess systematically different attitudes or characteristics than respondents, the resulting sample statistic is distorted.
  • Response Bias: Occurs when respondents provide inaccurate or untruthful answers due to social desirability pressures, fear of stigma, faulty memory, or intimidation by the interviewer.
  • Wording of Questions Bias: The use of leading, emotionally loaded, or confusing phrasing systematically steers respondents toward a specific viewpoint (e.g., asking 'Do you support investing in clean water to protect our children?' versus 'Do you support increasing municipal property taxes?').

The Four Pillars of Experimental Design

To isolate treatment effects and establish causal validity, an experiment must be structured around four core principles:

  1. Comparison: An experiment must compare two or more treatments (e.g., an experimental medication versus a control group) to ensure that observed outcomes are not simply the consequence of time trends, maturation, or external environmental fluctuations.
  2. Randomization: Experimental units must be assigned to treatment groups using an objective chance procedure. Random assignment neutralizes the effect of confounding variables by dispersing both known and unknown lurking variables evenly across all treatment conditions on average.
  3. Control: Researchers must keep extraneous variables constant across all treatment groups (e.g., maintaining identical temperature, lighting, and administration schedules) to ensure that the treatment is the sole variable operating. A control group provides an essential baseline, receiving either a standard existing treatment or a placebo (an inert dummy treatment). This isolates the placebo effect, wherein subjects exhibit measurable improvements purely due to their psychological expectation of receiving treatment.
  4. Replication: Treatments must be administered to a sufficiently large number of experimental units so that genuine treatment differences can be distinguished from random background variation. Replication does not mean repeating the entire study; rather, it refers to applying the treatment to multiple independent units within the experiment.

To prevent psychological bias, researchers employ single-blind designs (where subjects do not know which treatment they receive) or double-blind designs (where neither the participating subjects nor the clinicians and evaluators measuring the outcomes know who received which treatment). When a known extraneous variable (such as age, biological sex, or baseline blood pressure) is expected to influence the response, researchers utilize a randomized block design: subjects are first partitioned into homogeneous blocks, and treatments are then randomly assigned independently within each block.


Comparative Synthesis of Sampling Methodologies

Sampling MethodSelection MechanismGroup StructureKey AdvantagesPrimary Vulnerabilities / Risks
Simple Random Sample (SRS)Objective chance (random number generator, lottery)None; population treated as a single uniform groupUnbiased; mathematically straightforward; equal subset probabilityLogistically difficult for large, dispersed populations; high operational costs
Stratified Random SampleDraw an independent SRS from every single subgroupHomogeneous within strata, heterogeneous between strataGuarantees representation of all subgroups; reduces sampling varianceRequires prior knowledge of population stratification variables
Cluster SampleRandomly select entire subgroups; census all members within chosen unitsHeterogeneous within clusters, homogeneous between clustersMaximizes geographical and operational efficiency; lowers survey costIncreased sampling error if clusters differ substantially from one another
Systematic SampleSelect every $k$-th item from an ordered list after random startSequential ordering across populationEasy to execute in field work and continuous industrial processesSusceptible to extreme bias if list possesses hidden cyclical periodicity
Convenience SampleSelect individuals easiest to access or interviewUncontrolled opportunistic groupingQuick, inexpensive, minimal operational planning requiredExtreme selection bias; systematically unrepresentative; invalid inference
Voluntary ResponseIndividuals self-select by choosing whether to respondSelf-selected respondent groupInexpensive; rapid collection of large numbers of responsesSevere non-response and self-selection bias; overrepresents extreme views
Test Your Knowledge

A school district curriculum director evaluates an interactive digital mathematics platform. Twenty middle school algebra classrooms are randomly assigned to use the platform for one semester, while another twenty classrooms continue using the standard textbook curriculum. At the end of the semester, students in the digital platform classrooms achieve a statistically significantly higher mean score on a standardized algebra assessment. Which conclusion is statistically justified?

A
B
C
D
Test Your Knowledge

A public health researcher investigating community health needs in a large metropolitan county containing 80 distinct residential census tracts wants to gather representative survey data. The researcher assigns each tract a number from 1 to 80, uses a random number generator to select 8 entire census tracts, and surveys every household residing within those 8 chosen tracts. Which sampling design was executed?

A
B
C
D
Test Your Knowledge

A municipal transit authority places large posters on city buses directing commuters to an online survey that asks: 'To reduce dangerous carbon emissions and protect our environment, should the city council allocate additional funds to expand rapid bus transit lines?' Passengers scan a quick-response code to submit their votes. Of the 1,420 responses received, 88% vote in favor of expanding funding. Which two methodological vulnerabilities most severely compromise the validity of this survey as an estimate of general citywide public opinion?

A
B
C
D
Test Your Knowledge

An agronomist investigates the yield of a new drought-resistant maize hybrid compared to a standard variety. Soil moisture varies dramatically across the test farm from east to west due to a gentle slope. The agronomist divides the farm into four north-south strips of uniform soil moisture, divides each strip into two equal plots, and randomly assigns the drought-resistant hybrid to one plot and the standard variety to the other within each strip. What is the primary statistical justification for this randomized block design?

A
B
C
D