14.3 Threats to Internal and External Validity, Demand Characteristics, and Experimenter Bias

Key Takeaways

  • Campbell and Stanley's internal/external distinction, expanded by Cook and Campbell (1979), yields four core validity domains: internal validity (causal certainty), external validity (generalizability across settings and populations), construct validity (fidelity of operationalizations), and statistical conclusion validity (robustness of inferential tests).

  • Classic threats to internal validity include history, maturation, testing effects, instrumentation decay, statistical regression toward the mean, selection bias, differential attrition, and social threats such as compensatory rivalry (the John Henry effect).

  • Participant reactivity threatens validity through demand characteristics (cues revealing the hypothesis), evaluation apprehension, and the Hawthorne effect, where individuals alter behavior simply due to being observed.

  • Experimenter expectancy bias (the Rosenthal / Pygmalion effect) operates through subtle, unconscious cues that systematically guide participants toward confirming the researcher's hypothesis, as historically demonstrated by the Clever Hans phenomenon.

  • Double-blind randomized controlled trials (RCTs), active placebos, deception with ethical debriefing, and the Solomon Four-Group design provide methodological defenses against expectancy, testing artifacts, and confounding.

Last updated: October 2026

Threats to Internal and External Validity, Demand Characteristics, and Experimenter Bias

The ultimate goal of scientific experimentation is to produce inferences that are both internally defensible and externally applicable. An experiment can be executed with technical sophistication, yet its conclusions can be rendered completely uninterpretable by subtle confounds, participant expectations, or experimenter artifacts. In their landmark work, Donald Campbell and Julian Stanley (1963), along with Thomas Cook (1979), formalized the definitive taxonomy of validity threats in behavioral research. Understanding these threats and deploying rigorous methodological counter-measures is vital for evaluating empirical research on the GRE Psychology Subject Test.


1. The Campbell & Stanley Validity Taxonomy

Campbell and Stanley (1963) distinguished internal from external validity; Cook and Campbell (1979) added statistical conclusion validity and construct validity, yielding four interrelated conceptual domains:

                                  [ EXPERIMENTAL VALIDITY ]
                                              │
          ┌───────────────────────────┬───────┴───────────────────┬───────────────────────────┐
          ▼                           ▼                           ▼                           ▼
[ Internal Validity ]       [ External Validity ]       [ Construct Validity ]      [ Statistical Conclusion ]
Causal attribution to IV    Generalizability across     Operationalization fidelity   Proper statistical tests,
without confounds           populations & settings      of theoretical variables      power, & effect precision
  1. Internal Validity: The degree to which the observed changes in the dependent variable can be unambiguously and definitively attributed to the manipulation of the independent variable, rather than extraneous confounding factors. Internal validity is the core priority of basic experimental psychology.
  2. External Validity: The degree to which the causal relationships identified in an experiment can be legitimately generalized across different populations of people, operational variations of treatments, physical settings, and historical time points. Includes ecological validity—the extent to which experimental conditions mirror naturalistic, real-world psychological contexts.
  3. Construct Validity: The extent to which the empirical operationalizations of the independent and dependent variables accurately reflect the higher-order theoretical constructs they purport to instantiate. If an investigator operationalizes 'intelligence' solely as the speed of pressing a button, the study possesses poor construct validity.
  4. Statistical Conclusion Validity: The degree to which the researcher has applied appropriate statistical tests, satisfied parametric assumptions, maintained adequate statistical power (1−β≥.801 - \beta \ge .80), avoided inflated Type I error rates (from p-hacking), and generated accurate, reliable estimates of effect sizes (d,η2,rd, \eta^2, r).

The Fundamental Tension: Internal vs. External Validity

There is an inherent methodological trade-off between internal and external validity. To maximize internal validity, researchers isolate participants in artificial, tightly regulated laboratory environments, holding all environmental variables invariant. However, this artificiality often degrades external validity because human behavior in sterile laboratories may not generalize to chaotic, ecologically valid real-world settings. Conversely, field experiments conducted in naturalistic environments boast high ecological validity, but sacrifice internal validity due to uncontrolled extraneous variables.

2. Classic Threats to Internal Validity

Campbell and Stanley identified eight primary confounding threats that compromise internal validity, providing alternative explanations for observed changes between pretest and posttest:

                                [ THREATS TO INTERNAL VALIDITY ]
                                               │
  ┌───────────────┬───────────────┬────────────┴──┬───────────────┬───────────────┬───────────────┐
  ▼               ▼               ▼               ▼               ▼               ▼               ▼
History       Maturation       Testing      Instrumentation   Regression     Selection        Attrition
(External     (Internal       (Practice       (Observer /        to Mean      (Baseline     (Differential
 Events)      Biological)      Effects)       Tool Decay)     (Extremes)    Differences)       Drop-out)
Internal Validity ThreatOperational MechanismReal-World Experimental ExemplarMethodological Defense
HistorySpecific external environmental events occurring between the pretest and posttest that influence the DV, outside the experimenter's control.An economic depression or campus tragedy occurs during a semester-long intervention studying collegiate depression.Inclusion of an equivalent randomized control group exposed to the exact same historical timeline.
MaturationIntrinsic biological, physiological, or psychological changes occurring within participants purely as a function of the passage of time (e.g., aging, fatigue, spontaneous remission).In evaluating a reading intervention for 6-year-olds over a year, children improve naturally due to biological brain development and baseline schooling.Randomized control group; comparison against normative maturational trajectories.
Testing (Practice)The psychological or cognitive impact of taking a pretest upon the scores of a subsequent posttest, independent of any experimental treatment.Participants taking an IQ test at pretest score higher at posttest simply because they are familiar with the item formats and test-taking strategies.Omission of pretest (posttest-only design); use of the Solomon Four-Group Design.
Instrumentation (Decay)Changes in the measuring instrument, observational coding standards, or human raters across the duration of the experiment.Observers become more experienced (or fatigued and careless) over time, systematically altering how they rate childhood aggression between week 1 and week 12.Rigorous rater training, objective automated data collection, evaluating inter-rater reliability (Cohen′s κCohen's\ \kappa).
Statistical Regression toward the MeanWhen participants are selected for an intervention based on extreme baseline scores (extremely high or extremely low), their retest scores naturally gravitate toward the population mean due to measurement error.Depressed patients selected when their scores are at an absolute crisis peak show lower depression scores three months later purely due to statistical regression.Selecting participants across the full continuum; randomized control group to establish baseline regression drift.
Selection BiasNon-random assignment produces systematic, pre-existing baseline differences in participant characteristics between experimental and control groups.Comparing voluntary participants in an optional after-school tutoring program to non-attendees (volunteers possess higher baseline motivation).True random assignment of participants to experimental conditions.
Mortality (Differential Attrition)Non-random, systematic drop-out of participants from conditions during the study, altering the composition and equivalence of the groups.In an intense exercise trial, participants experiencing severe fatigue or no results drop out of the treatment group, leaving only highly fit, resilient outliers.Tracking baseline traits of dropouts vs. completers; intention-to-treat (ITT) statistical analysis.
Selection InteractionsSelection bias combines with another threat, such that the threat operates differently across groups (e.g., Selection-Maturation).In a study of head start education, children in the lower-SES control group mature cognitively at a different rate than higher-SES comparison children.True random assignment from a homogeneous participant pool.

3. Social and Treatment Diffusion Threats to Internal Validity

When human participants and research staff interact within an organizational or community setting, social processes can generate unique confounds that undermine experimental integrity:

                          [ SOCIAL THREATS TO INTERNAL VALIDITY ]
                                             │
          ┌──────────────────────────┬───────┴──────────────────┬──────────────────────────┐
          ▼                          ▼                          ▼                          ▼
[ Diffusion of Treatment ] [ Compensatory Rivalry ]   [ Compensatory Equalization ] [ Resentful Demoralization ]
Control group adopts       Control group works        Administrators provide extra  Control group gives up,
experimental techniques    harder to compete          resources to control group    exhibiting artificially low
via cross-communication    ('John Henry Effect')      to compensate for deprivation performance
  • Diffusion of Treatment: Occurs when participants in the experimental group communicate with, or share intervention materials with, participants in the control group. The control group inadvertently adopts the active ingredients of the treatment, reducing between-group differences and causing the researcher to falsely conclude that the intervention had no effect.
  • Compensatory Rivalry (The John Henry Effect): When participants in the control group discover that they are in the untreated comparison condition, they may perceive themselves as being at a competitive disadvantage. Motivated to prove their worth, they exert extraordinary, uncharacteristic compensatory effort to outperform the experimental group. This artificial elevation of control performance masks genuine treatment effects.
  • Compensatory Equalization of Treatment: Occurs when research administrators, clinicians, or teachers feel sympathetic toward the untreated control group because they are being deprived of a potentially beneficial intervention. To compensate, staff provide alternative goods, extra attention, or supplemental services to the control group, contaminating the baseline.
  • Resentful Demoralization: The opposite of compensatory rivalry: control group participants become discouraged, resentful, or angry upon learning that they were denied a desirable or prestigious treatment. Consequently, they withdraw effort, exhibit apathy, or perform artificially worse, exaggerating the apparent superiority of the experimental treatment.

4. Participant and Experimenter Biases

In psychological science, the human beings being observed—and the human beings conducting the observation—introduce systematic cognitive and social biases that can severely distort empirical data.

Participant Biases and Reactivity

  • Demand Characteristics: First systematically characterized by Martin Orne (1962), demand characteristics are subtle, explicit, or implicit environmental cues within an experimental setting that convey the researcher's true hypothesis to the participant. Participants are rarely passive responders; they actively construct theories about the study and adopt specific participant roles:
    • The Good Subject: Eagerly alters behavior to confirm the experimenter's perceived hypothesis.
    • The Bad (Negativistic) Subject: Intentionally acts to disrupt, contradict, or invalidate the hypothesis.
    • The Apprehensive Subject: Anxious about being evaluated; behaves unnaturally to appear socially desirable, mentally healthy, or highly competent (Evaluation Apprehension, Rosenberg, 1969).
  • The Hawthorne Effect: Originating from industrial efficiency studies at Western Electric's Hawthorne Works (1924–1932). In the early illumination tests (1924–1927), company and National Research Council engineers manipulated factory lighting to see if brighter illumination increased worker productivity; Elton Mayo and Fritz Roethlisberger later led the relay-assembly and interviewing phases and popularized the findings. Productivity increased under bright light—but it also increased when lights were dimmed. The researchers realized that individuals alter their behavior simply because they are aware that they are being observed and evaluated, regardless of the physical experimental manipulation. (Later reanalyses of the original records suggest the effect was smaller and less consistent than the classic story implies.)
  • The Placebo Effect: A genuine psychological or neurobiological improvement (e.g., pain reduction, alleviation of depressive symptoms) produced not by an active therapeutic agent, but by the participant's psychological expectation and belief that they are receiving an efficacious treatment. Placebo analgesia, for example, is mediated by genuine endogenous opioid release in the brain and can be blocked by the opioid antagonist naloxone.

Experimenter Biases

  • Experimenter Expectancy Bias (The Rosenthal Effect): Documented extensively by Robert Rosenthal, this phenomenon occurs when an experimenter's preconceived hypotheses or expectations unconsciously lead them to treat participants in subtle, systematic ways that steer participants toward confirming the hypothesis (e.g., subtle changes in tone of voice, posture, smiling, nodding, or selective recording of ambiguous responses).
  • The Pygmalion Effect (Rosenthal & Jacobson, 1968): In Pygmalion in the Classroom, teachers were falsely informed that certain randomly selected elementary students were 'academic bloomers' who would show intellectual spurts. At the end of the year, those randomly labeled students exhibited significantly greater objective gains in IQ scores than control children. Teacher expectancies unconsciously altered pedagogical warmth, feedback quality, and cognitive challenge, transforming expectation into reality.
  • The Clever Hans Effect: In early 20th-century Germany, a horse named Clever Hans appeared to solve complex arithmetic calculations, spell words, and tell time by tapping his hoof. In 1907, psychologist Oskar Pfungst rigorously investigated Hans and revealed that the horse possessed zero mathematical aptitude. Instead, Hans was exquisitely sensitive to the unconscious, involuntary physical micro-cues of his human questioners (e.g., slight head tilts, changes in breathing, muscular tension) who leaned forward to watch his hoof and straightened up minutely when the correct count was reached.

5. Methodological Controls for Validity Threats and Biases

To safeguard the integrity of psychological findings, researchers implement specialized structural control designs:

Single-Blind RCT:    [ Participant Blinded ] ──► Controls for Demand Characteristics & Placebo Effects
Double-Blind RCT:    [ Participant & Experimenter Blinded ] ──► Controls for Demand, Placebo, & Expectancy
Active Placebo:      [ Mimics Side Effects without Active Drug ] ──► Preserves Blinding Integrity

Blinding Protocols

  • Single-Blind Study: Participants are kept naive regarding which experimental condition they have been assigned to, preventing demand characteristics and placebo expectations.
  • Double-Blind Randomized Controlled Trial (RCT): Neither the research participant nor the experimenter directly interacting with them (or scoring the dependent outcome) knows which treatment condition the participant has been assigned to. An independent pharmacy or third-party administrator manages condition coding. Double-blind RCTs represent the gold standard in psychopharmacology and clinical trials because they simultaneously eliminate participant placebo effects and experimenter expectancy biases.
  • Active Placebos: In psychopharmacological trials, inert sugar pills fail to maintain blinding if the experimental drug produces noticeable autonomic side effects (e.g., dry mouth, dizziness, nausea). Participants deduce they are in the active group, restoring expectancy bias. An active placebo contains an inert substance combined with a mild agent (such as atropine) that mimics the somatic side effects of the experimental drug without possessing its therapeutic neurochemical action, successfully preserving double-blind integrity.

Deception and Ethical Debriefing

To eliminate demand characteristics in social and cognitive psychology, researchers frequently employ deception by providing a plausible cover story that obscures the true hypothesis. APA ethical standards permit deception only when: (1) no viable non-deceptive alternative exists, (2) the study possesses significant scientific or applied value, and (3) participants receive immediate, comprehensive debriefing upon study completion. Debriefing must include dehoaxing (revealing the true hypothesis and experimental deceptions) and desensitizing (removing any distress or negative affect induced by the study).

The Solomon Four-Group Design

The Solomon Four-Group Design is the definitive experimental architecture for assessing and eliminating the confounding effects of pretesting:

Group 1 (Experimental with Pretest):   Pretest (O1) ──► Treatment (X) ──► Posttest (O2)
Group 2 (Control with Pretest):        Pretest (O3) ──► Control (---)  ──► Posttest (O4)
Group 3 (Experimental without Pretest):                 Treatment (X) ──► Posttest (O5)
Group 4 (Control without Pretest):                      Control (---)  ──► Posttest (O6)

By comparing posttest scores across these four groups, researchers can execute three critical diagnostic comparisons:

  1. Evaluate the Main Effect of the Treatment: Comparing (O2+O5)(O_2 + O_5) against (O4+O6)(O_4 + O_6) determines whether the treatment (X)(X) exerted a genuine causal effect independent of pretesting.
  2. Evaluate the Main Effect of Pretesting: Comparing (O2+O4)(O_2 + O_4) against (O5+O6)(O_5 + O_6) reveals whether the simple act of taking a pretest altered posttest scores.
  3. Detect Pretest-by-Treatment Interaction: Evaluates whether receiving a pretest sensitizes participants, causing them to respond to the treatment differently than unpretested individuals.

Note

A favorite GRE Subject Test question presents an experimental scenario where an investigator wants to determine whether an educational video enhances learning, but worries that administering a baseline pretest will prime participants to look for specific answers in the video. The definitive methodological solution to isolate this pretest sensitization is the Solomon Four-Group Design.

Test Your Knowledge

A clinical psychologist evaluates a new 8-week mindfulness intervention for severe major depressive disorder. The researcher recruits twenty patients whose Beck Depression Inventory scores are in the top 2% of the clinical population. At the conclusion of the eight weeks, the patients show a statistically significant reduction in depressive symptoms. Why can the researcher NOT validly attribute this improvement to the mindfulness intervention?

A

The Hawthorne effect caused the patients to perform worse due to being observed by clinical staff.

B

The study suffered from differential participant attrition that eliminated the lowest-functioning patients before posttest.

C

The experimenter inadvertently created an uncalibrated ceiling effect on the depression inventory.

D

Statistical regression toward the mean, because extreme baseline scores tend to drift toward the average on retesting.

Test Your Knowledge

In a classic study evaluating cognitive expectancies, Robert Rosenthal and Lenore Jacobson informed elementary school teachers that certain students were identified by a psychological test as 'intellectual bloomers,' even though the children were actually selected at random. By the end of the academic year, these randomly selected children showed significant, objective gains in IQ scores. What phenomenon does this finding demonstrate?

A

The John Henry effect of compensatory rivalry

B

Evaluation apprehension driven by demand characteristics

C

The Pygmalion effect of experimenter and teacher expectancy

D

Instrument decay resulting from shifting standardized testing criteria

Test Your Knowledge

An experimental psychologist is concerned that administering a baseline pretest on memory strategies will prime participants to attend to specific mnemonic techniques during an upcoming training video, altering how they respond to the treatment. Which of the following experimental designs is specifically constructed to detect and isolate this pretest sensitization effect?

A

A non-equivalent control group pretest-posttest design

B

A balanced Latin Square repeated-measures design

C

The Solomon Four-Group Design

D

A single-blind regression-discontinuity design

Test Your Knowledge

During a multi-site randomized trial comparing a novel computerized reading curriculum against a standard curriculum, teachers in the control classrooms learn that their students are being compared to the experimental classrooms. Feeling competitive, the control teachers voluntarily spend extra hours providing individualized phonics instruction to ensure their students are not outperformed. What specific threat to internal validity has occurred?

A

Compensatory rivalry (The John Henry Effect)

B

Diffusion of treatment

C

Resentful demoralization of the control teachers

D

Statistical regression toward the mean

Sections you finish are checked off in the contents.