11.1 Statistical Studies: Surveys, Experiments, and Observational Studies
Key Takeaways
- A sample survey gathers data from a representative subset of a population using random selection to generalize sample statistics to population parameters without imposing treatments.
- Observational studies record variables without intervention, identifying correlations or associations between variables while being fundamentally incapable of establishing cause-and-effect relationships due to potential confounding variables.
- A randomized controlled experiment randomly assigns subjects to treatment and control conditions; it is the only study design capable of establishing a causal relationship between variables.
- Core experimental controls include comparison groups, placebos to account for the psychological placebo effect, and blinding protocols (single-blind and double-blind) to eliminate participant and researcher bias.
- Valid statistical inference requires probability-based sampling (simple random, stratified, cluster) to prevent selection bias, whereas convenience and voluntary response sampling introduce severe systematic bias.
11.1 Statistical Studies: Surveys, Experiments, and Observational Studies
Quick Answer: The three main types of statistical inquiries are sample surveys, observational studies, and randomized controlled experiments. Only a randomized controlled experiment—in which researchers actively impose treatments through random assignment—can establish a cause-and-effect relationship. Observational studies and surveys can detect associations or correlations, but confounding variables prevent causal claims. Furthermore, random selection of subjects allows findings to be generalized to a larger population, whereas random assignment of treatments isolates the causal impact of the explanatory variable.
1. Taxonomy of Statistical Inquiries (AII-S.IC.3, AII-S.IC.6b)
[!NOTE] Standard note. AII-S.IC.3 is the assessed standard - recognize the purposes of and differences among sample surveys, experiments, and observational studies; explain how randomization relates to each. NYSED removed S-IC.1 as a separate standard, noting that the expectation persists inside the other S-IC standards. Critiquing a study\u2019s claims in ordinary language belongs to AII-S.IC.6b (use the language of statistics to critique claims from informational texts - for example, causation versus correlation, bias, measures of center and spread). The June 2026 Part I asked students to decide whether a driving-simulator study was observational, given that researchers randomly assigned the two conditions.
In statistical inference, the methodology used to collect data dictates the mathematical scope and validity of any resulting conclusions. Regents examination questions frequently require students to classify a study design and determine whether a proposed causal claim or generalization is scientifically justified.
Statistical Studies
│
┌─────────────────────────────────┴─────────────────────────────────┐
▼ ▼
Observational Experimental
(No Treatments) (Treatments Imposed)
│ │
┌─────┴──────────────┐ │
▼ ▼ ▼
Sample Survey Observational Study Randomized Experiment
(Generalize to (Identify Associations; (Random Assignment;
Population) Confounding Present) Establishes Causation)
Sample Surveys
A sample survey collects data from a subset (sample) of a defined population without manipulating or modifying the environment. The primary objective is to estimate unknown population parameters (such as the proportion of voters favoring a policy or the mean household income in a county).
- Mechanism: Researchers measure characteristics as they naturally exist.
- Scope of Inference: When the sample is selected using a probability-based random sampling technique, the sample statistics can be generalized to the entire target population from which the sample was drawn.
- Limitation: Surveys measure opinions or self-reported behaviors; they do not establish causal mechanisms.
Observational Studies
In an observational study, researchers observe and measure variables of interest without assigning treatments or influencing the subjects. For example, a researcher tracking dietary habits and cardiovascular health over ten years simply records what participants eat and monitors health outcomes.
- Value: Essential for studying conditions where experimental manipulation is unethical, dangerous, or physically impossible (e.g., studying the health consequences of smoking, exposure to environmental toxins, or early childhood trauma).
- Fatal Inferential Flaw: Observational studies cannot establish cause-and-effect. Even if an overwhelming association is discovered between two variables, confounding variables (lurking external factors) may be the true driver of the observed outcome.
Randomized Controlled Experiments
A randomized controlled experiment is a study design where investigators deliberately impose one or more treatments on experimental units (human subjects, animals, or objects) and observe the resulting responses. Crucially, the researchers use random assignment to allocate subjects among different treatment conditions.
- The Causality Benchmark: A well-designed, randomized controlled experiment is the only study design that allows researchers to conclude that the treatment caused the change in the response variable.
- Mechanism: Randomly assigning subjects distributes both known and unknown confounding variables evenly across treatment and control groups at the start of the study, isolating the explanatory variable as the sole systematic difference between groups.
Comparison of Study Designs
| Study Design | Treatment Imposed? | Random Mechanism Employed | Scope of Inference | Establishes Causation? |
|---|---|---|---|---|
| Sample Survey | No | Random Selection of Subjects | Generalizes to target population | No |
| Observational Study | No | None (Passive observation) | Identifies correlation/association only | No |
| Randomized Experiment | Yes | Random Assignment to Treatments | Establishes causal relationship | Yes |
2. The Golden Rule: Random Selection vs. Random Assignment
A central conceptual distinction on the Regents exam is the difference between random selection and random assignment. These two random mechanisms serve completely distinct inferential purposes:
+-------------------------------------------------------------+
| The 2 × 2 Scope of Inference Grid |
+-------------------------------------------------------------+
| Random Assignment to Treatments? |
| YES NO |
+-------------------------------------------------------------+
Random Selection | YES | Causation established; Generalize to |
from Population? | | Generalize to population population; No cause |
| NO | Causation established; No causation; |
| | Group participants only Anecdotal / Associations|
+-------------------------------------------------------------+
- Random Selection $\implies$ Generalizability: If subjects are randomly selected from a broader population, the sample is representative of that population, allowing conclusions to be generalized.
- Random Assignment $\implies$ Causality: If subjects are randomly assigned to treatment groups, confounding variables are balanced across groups, allowing differences in the response variable to be attributed directly to the treatment.
[!IMPORTANT] Regents Rule of Thumb: If an exam question asks, "Can the researchers conclude that the new reading program caused the increase in reading scores?" examine the study design. If subjects volunteered and chose their own method, or if entire existing classrooms were compared without random assignment of individual students, you must answer No, because confounding variables (such as student motivation or teacher effectiveness) were not controlled via random assignment.
3. Experimental Design Principles & Controlling Confounding
To ensure internal validity, experiments adhere to four foundational principles:
- Comparison / Control: An experiment must compare two or more treatments (e.g., new drug vs. existing standard drug, or new drug vs. placebo). The control group provides a baseline measurement against which the treatment effect is measured.
- Random Assignment: Experimental units must be assigned to treatment groups using a chance mechanism (e.g., random number generator, slip drawing, or coin toss). This balances lurking variables.
- Replication: Treatments must be applied to a sufficiently large number of subjects so that differences between groups can be distinguished from chance variation.
- Control of Variability / Blinding: Keeping all other environmental factors identical across groups prevents extraneous factors from influencing the response.
Placebos and the Placebo Effect
The placebo effect is the measurable psychological or physiological improvement in health or behavior resulting merely from the participant's belief that they are receiving an active treatment. A placebo is an inert, harmless substance (such as a sterile saline injection or starch pill) designed to look, taste, and smell identical to the experimental treatment. Incorporating a placebo group isolates the chemical efficacy of the medication from the psychological benefit of receiving medical care.
Blinding Protocols
To prevent psychological bias from contaminating experimental outcomes, researchers implement blinding:
- Single-Blind Study: The subjects do not know which treatment they are receiving (active drug or placebo), but the evaluating researchers know.
- Double-Blind Study: Neither the human subjects nor the diagnosing physicians/evaluators who measure the response variable know which treatment each subject received. Double-blinding prevents experimenter expectancy bias, where a doctor might unconsciously rate a patient's recovery more favorably if they know the patient received the active medication.
Confounding Variables
A confounding variable is an extraneous variable related to both the explanatory variable (treatment) and the response variable, in such a way that its specific effect on the response cannot be separated from the effect of the explanatory variable. For example, in an observational study showing that people who drink more coffee have higher rates of heart disease, cigarette smoking is a classic confounding variable: coffee drinkers historically smoked cigarettes at higher rates, and smoking directly elevates heart disease risk.
4. Sampling Techniques: Probability vs. Non-Probability
When conducting a sample survey, how the sample is chosen determines whether the data are scientifically meaningful.
Probability-Based Sampling (Statistically Valid)
- Simple Random Sample (SRS): Every individual in the population has an equal probability of selection, and every possible subset of size $n$ has an equal chance of being selected (e.g., generating $n$ distinct random integers from a student ID directory).
- Stratified Random Sample: The population is divided into non-overlapping, homogeneous subgroups called strata based on a shared characteristic (e.g., freshmen, sophomores, juniors, seniors). A simple random sample is then drawn independently from every stratum. This guarantees representation of all key demographic segments.
- Cluster Sample: The population is divided into naturally occurring, heterogeneous subgroups called clusters (e.g., homeroom classrooms or geographic neighborhoods). A random sample of entire clusters is selected, and every individual within the chosen clusters is surveyed.
- Systematic Sample: Subjects are selected at regular numerical intervals from an ordered list (e.g., surveying every $10^{\text{th}}$ person entering a building after a random starting point between 1 and 10).
Flawed / Non-Probability Sampling (Systematically Biased)
- Convenience Sampling: Researchers select individuals who are easiest to reach (e.g., asking friends in the cafeteria or surveying people walking past a shopping mall entrance). This systematically over-represents individuals present at that location and time.
- Voluntary Response Sampling: An open invitation is extended (e.g., online polls, call-in radio shows, social media votes), and participants choose whether to respond. This method suffers from extreme bias because people with strong negative or emotional opinions are far more likely to respond than the general public.
5. Systemic Sources of Bias in Statistical Studies
Bias is the systematic tendency of a study's design to favor certain outcomes over others, producing sample statistics that consistently overestimate or underestimate the true population parameter.
| Type of Bias | Mechanism | Real-World Example |
|---|---|---|
| Selection Bias (Undercoverage) | Certain groups in the population are systematically omitted or underrepresented in the sampling frame. | Conducting an election survey strictly via landline telephones, excluding younger voters who use only mobile phones. |
| Nonresponse Bias | Selected individuals cannot be contacted or refuse to participate, and nonrespondents differ systematically from respondents. | Mailing 1,000 surveys regarding work-life balance where 80% of recipients never return the form due to long working hours. |
| Response Bias | Participants provide inaccurate or untruthful answers due to question wording, interviewer behavior, or social stigma. | A survey asking, "Do you agree that our wonderful mayor deserves another term?" (leading question) or asking students in front of teachers if they cheat. |
6. Worked Problems
Worked Problem 1: Evaluating Study Design and Permissible Claims
Problem: A school board wants to investigate whether an after-school peer tutoring program improves scores on the Regents Algebra II exam. At West High School, 40 students voluntarily enroll in the tutoring program, while 120 students choose not to enroll. At the end of the semester, the mean Regents exam score for the tutored group is $84$, while the mean score for the untutored group is $76$. The board issues a press release stating: "Peer tutoring causes students to score an average of 8 points higher on the Regents exam."
- Identify whether this study was a sample survey, an observational study, or a randomized controlled experiment.
- Explain why the board's causal claim is statistically invalid, citing a specific confounding variable.
-
Step 1: Identify the study design. Although a comparison is made between two groups, students voluntarily chose whether to attend tutoring; researchers did not randomly assign students to tutoring or no tutoring. Therefore, this is an observational study.
-
Step 2: Evaluate the causal claim. Causal conclusions cannot be drawn from an observational study because treatments were not randomly assigned to subjects.
-
Step 3: Identify a plausible confounding variable. Student academic motivation is a critical confounding variable. Students who voluntarily attend after-school tutoring are likely already more motivated, study more hours independently, or have better attendance. This pre-existing motivation could be the true cause of the higher exam scores, rather than the tutoring program itself.
Worked Problem 2: Analyzing Sampling Techniques
Problem: A district superintendent wishes to gauge teacher satisfaction across 20 elementary schools containing 800 total teachers. Two sampling plans are proposed:
- Plan A: Post a feedback questionnaire on the district employee portal and collect the first 100 responses.
- Plan B: Organize teachers by school building (20 strata), randomly select 5 teachers from each building using a random number generator, and survey all 100 selected teachers.
Classify each sampling plan and determine which plan provides a representative sample with minimal bias.
-
Step 1: Classify Plan A. Plan A relies on voluntary participation and convenience (taking the first 100 respondents). This is a voluntary response / convenience sample and is prone to severe nonresponse bias and selection bias, as dissatisfied teachers or those with extra free time may respond first.
-
Step 2: Classify Plan B. Plan B partitions the population into homogeneous geographic strata (individual schools) and draws an independent simple random sample of equal size ($5$ teachers) from each stratum. This is a stratified random sample.
-
Step 3: Conclude which plan minimizes bias. Plan B is statistically sound because random selection ensures every teacher in every building has an equal probability of selection, guaranteeing balanced geographical representation across the district and minimizing selection bias.
7. Common Regents Pitfalls & Exam Strategies
- Pitfall 1: Conflating Random Sampling with Random Assignment. Random sampling determines who is in the study (governing generalizability to the population). Random assignment determines which treatment each subject receives (governing cause-and-effect conclusions). An experiment can establish causation without having a random sample of the world population.
- Pitfall 2: Overlooking Blindness in Subjective Evaluations. On free-response questions asking why a double-blind design is necessary for a pain-relief trial, always state: "To prevent the researchers from unconsciously rating patient outcomes differently based on knowledge of which treatment was administered."
- Pitfall 3: Claiming a Large Voluntary Sample Eliminates Bias. A voluntary poll of 50,000 internet users is still fundamentally biased; increasing sample size in a biased sampling design merely produces a more precise estimate of a biased number.
A botanist wants to determine whether a newly formulated mineral spray increases the yield of commercial tomato plants. She randomly divides 60 identical tomato seedlings into two groups of 30: one group receives the mineral spray weekly, while the other group receives an identical volume of plain water. All plants are grown in the same greenhouse with identical soil, sunlight, and temperature. After 10 weeks, the sprayed plants produce a mean yield of 5.1 kg per plant, while the unsprayed plants produce a mean yield of 3.4 kg per plant. Which statement represents a statistically valid conclusion?
The principal of a suburban high school with 1,500 students wants to assess student satisfaction with the school cafeteria's menu. An optional survey link is posted on the school's social media page, and 92 students complete the questionnaire. Which statement best identifies the primary statistical flaw in this study's design?
A medical research institute evaluates a new allergy medication. Two hundred participants are randomly assigned to receive either the active antihistamine tablet or an identical-looking sugar tablet. Neither the participants nor the examining allergists know which tablet each subject received until all symptom relief scores are recorded. Which combination of experimental design features is present in this study?