1.3 Evaluating Experimental Design, Sample Size, Validity & Sources of Error

Key Takeaways

  • Internal validity requires isolating the independent variable so that observed changes in the dependent variable can be attributed solely to the experimental manipulation.
  • External validity assesses whether experimental findings can be generalized beyond the laboratory setting to real-world populations or ecosystems.
  • Large sample sizes and multiple experimental trials reduce the influence of random error and individual anomalies, yielding representative mean values.
  • Systematic errors consistently skew data in one direction due to equipment or procedural defects and cannot be corrected by averaging repeated trials, whereas random errors cause unpredictable fluctuations around the true mean that cancel out with repetition.
Last updated: September 2026

Evaluating Experimental Design, Sample Size, Validity & Sources of Error

Quick Answer: A valid experiment tests only one independent variable at a time, uses an adequate sample size, includes control groups, and holds all other factors constant. Internal validity means the experiment truly measured what it intended to measure without confounding variables. Systematic errors (such as an uncalibrated scale) skew all data in one direction and cannot be fixed by averaging trials, while random errors (unpredictable fluctuations) cancel out through large sample sizes and repeated trials.

On the HiSET Science test, several questions present experimental setups that contain subtle or obvious flaws. Your task is to evaluate whether the methodology supports the stated conclusions, identify uncontrolled variables, and recommend procedural improvements.


Methodological Rigor: Validity and Reliability

High-quality scientific research rests on three core standards:

1. Internal Validity

Internal validity measures how conclusively an experiment proves that changes in the dependent variable were caused solely by the independent variable. If an unmonitored factor (such as temperature or soil type) varies between groups, it becomes a confounding variable, destroying internal validity and preventing causal claims.

2. External Validity (Generalizability)

External validity reflects how well laboratory findings apply to real-world populations or natural ecosystems. A fertilizer that doubles yield in a sterile greenhouse might fail in open fields subject to drought and pests. While the greenhouse study has high internal validity, external validity requires field testing.

3. Reliability and Reproducibility

  • Reliability (Precision): The degree to which a procedure yields consistent measurements under identical conditions.
  • Reproducibility: The ability of independent researchers to duplicate the experiment following published methods and obtain matching results. Findings that cannot be replicated are rejected.

Confounding Variables and Methodological Flaws

A confounding variable is an uncontrolled extraneous factor that correlates with both independent and dependent variables, obscuring the true relationship. Common design flaws tested on the HiSET include:

  • Failure to Standardize Conditions: Exposing test groups to differing parameters (e.g., sunlight vs. shade).
  • Selection Bias: Non-random assignment of subjects. Placing healthier subjects in the treatment group biases results.
  • Lack of Placebo / Blinding Controls: In clinical studies, participant expectations can trigger changes (the placebo effect). To eliminate bias:
    • Single-Blind Study: Participants do not know whether they are receiving the active treatment or a placebo.
    • Double-Blind Study: Neither participants nor researchers administering treatments know group assignments until data collection ends.

Sample Size, Replication, and Statistical Power

A frequent flaw in flawed experiments is an insufficient sample size (n). Testing only one or two subjects yields unreliable data:

The Danger of Small Sample Sizes

Living organisms exhibit natural genetic and environmental variation. If a hormone is tested on two animals, one may grow rapidly due to genetics while the other suffers an unobserved illness. With n = 2, individual anomalies dominate data, producing false conclusions.

The Law of Large Numbers

As sample size increases (n ≥ 30, 100), individual anomalies and measurement noise average out. The sample mean approaches the true population mean, and standard error decreases, giving the statistical power needed to detect genuine treatment effects.

Repeated Trials (Replication)

Every experiment must incorporate multiple independent trials (k ≥ 3). Testing only once leaves results vulnerable to transient equipment fluctuations or contamination. Replication confirms consistency, allowing researchers to calculate standard deviation and identify outliers.


Categorizing Scientific Error: Systematic vs. Random Error

Scientists distinguish between two fundamentally different categories of experimental error:

FeatureSystematic Error (Bias)Random Error (Noise)
DefinitionPredictable bias that skews all measurements consistently in one direction (+ or -).Unpredictable fluctuations that cause measurements to scatter above and below the true value.
Primary CausesMiscalibrated instruments (e.g., scale reading +0.5 g too high), zero-offset errors, contaminated reagents.Minor drafts, electrical noise, human eye parallax when reading a meniscus, slight temperature shifts.
Impact on DataDestroys accuracy (closeness to true value) while maintaining apparent precision.Decreases precision (repeatability) while the average value may still center on the true value.
Reduced by Averaging?NO. Repeating the measurement 1,000 times reproduces the identical skewed bias every time.YES. Because random fluctuations occur equally in positive and negative directions, averaging cancels them out.
Corrective ActionRecalibrate instruments, redesign protocols, re-zero balances.Increase sample size (n) and collect multiple repeated measurements.

HiSET Scenario Walkthrough: Auditing a Flawed Experiment

Flawed Scenario: A student tests whether Pesticide X protects tomato plants from aphids better than water alone. She uses 6 plants:

  • Group 1 (Pesticide X): 3 large, flowering plants in loam soil in a greenhouse are sprayed with Pesticide X.
  • Group 2 (Control): 3 small seedlings in sandy soil in an outdoor garden are sprayed with water.

After 4 weeks, Group 1 has fewer aphids and higher fruit yield. The student concludes Pesticide X is highly effective.

Methodological Audit & Critique

This experiment contains severe flaws that invalidate the conclusion:

  1. Confounding Variables: Plant maturity, soil type, moisture, sunlight, and temperature all differed between groups.
  2. Inadequate Sample Size: Testing only 3 plants per group provides virtually no statistical power; a single resistant plant could skew results.
  3. Selection Bias: Healthier plants were placed in the treatment group.
  4. Redesign Protocol: Use at least 50 uniform seedlings of identical age; pot them in identical soil; keep all plants in the same greenhouse; randomly assign 25 to Pesticide X and 25 to water; record aphid counts at regular intervals.
Loading diagram...
Methodological Audit: Evaluating Scientific Validity and Reliability
Test Your Knowledge

A researcher conducts an experiment to test whether adding crushed basalt rock to agricultural soil accelerates plant carbon capture. Group A consists of 4 mature corn plants grown in clay soil in an open field, receiving crushed basalt. Group B consists of 4 immature soybean plants grown in sandy soil inside a climate-controlled greenhouse, receiving no basalt. What is the primary methodological flaw that prevents the researcher from drawing a valid scientific conclusion?

A
B
C
D
Test Your Knowledge

An analytical digital balance in a chemistry laboratory is improperly calibrated, consistently recording masses that are exactly 0.50 grams heavier than the true mass of any substance placed on the pan. Which of the following correctly describes this type of error and its remediation?

A
B
C
D
Test Your Knowledge

In a randomized clinical trial evaluating the therapeutic efficacy of a new cholesterol-lowering medication, researchers implement a double-blind protocol in which neither the patients nor the physicians administering the pills know who receives the active compound and who receives an identical inert sugar pill. What is the fundamental scientific rationale for using a double-blind design?

A
B
C
D