13.2 Validity, Reliability & Confounding Variables
Key Takeaways
- Internal validity measures how confidently changes in the dependent variable can be attributed to the independent variable, whereas external validity measures generalizability to broader populations and real-world settings.
- Reliability refers to consistency and precision of measurement; reliability is a necessary, but not sufficient, condition for validity.
- Confounding variables create spurious relationships by independently influencing both the IV and DV, whereas mediating variables explain the causal mechanism (HOW/WHY), and moderating variables alter the strength of the effect (WHEN/FOR WHOM).
- The Hawthorne effect occurs when participants alter their behavior due to awareness of being observed, threatening external and internal validity.
- Experimenter bias can be neutralized using double-blind protocols, while demand characteristics are minimized through covert or unobtrusive measures.
13.2 Validity, Reliability & Confounding Variables
In MCAT scientific reasoning, evaluating the rigor of experimental and observational research requires analyzing measurement precision, internal control, external generalizability, and potential sources of systematic error.
Internal vs. External Validity
Scientific validity refers to the correctness or truthfulness of inferences drawn from a study. Psych/Soc passages frequently test the tension between internal validity and external validity.
HIGH INTERNAL VALIDITY <===================> HIGH EXTERNAL VALIDITY
(Tight Lab Controls, Artificial Environment) (Real-World Realism, Naturalistic Sample)
Internal Validity
Internal Validity is the degree to which an experiment demonstrates that changes in the independent variable (IV) directly caused the observed changes in the dependent variable (DV), free from alternative explanations or confounding influences.
Primary Threats to Internal Validity:
- Confounding Variables: Uncontrolled extraneous factors that correlate with both the IV and DV, masking or artificially creating relationships.
- Selection Bias: Non-random assignment resulting in baseline differences between experimental and control groups.
- Attrition / Mortality Bias: Non-random loss of participants over time (e.g., sicker patients dropping out of an arduous drug trial), leaving an unrepresentative residual sample.
- Maturation: Natural biological or psychological changes occurring within participants over time (e.g., children growing older, spontaneous recovery from illness) independent of the intervention.
- History Effects: External environmental events occurring between pre-test and post-test measurements that influence the DV (e.g., a national economic crisis occurring during a study on job stress).
- Testing Effects / Practice Effects: Changes in DV scores resulting from prior exposure to the testing instrument itself.
- Instrumentation Changes: Unintended alterations in measurement instruments, observer rating criteria, or calibration over the course of a study.
- Regression to the Mean: The statistical phenomenon wherein extreme scores on an initial assessment naturally move closer to the population mean upon re-testing, independent of any intervention.
External Validity
External Validity is the extent to which study findings can be generalized beyond the specific experimental sample to other populations, settings, times, and operational conditions.
Primary Threats to External Validity:
- Sampling Bias / Non-Representative Samples: Selecting participants from narrow demographics (e.g., conducting psychological studies exclusively on WEIRD populations—Western, Educated, Industrialized, Rich, and Democratic college undergraduates).
- Artificiality of Experimental Setting (Low Ecological Validity): Highly artificial laboratory environments that fail to mimic real-world physical, social, or emotional contexts.
- Hawthorne Effect (Observer Effect): Participants alter their natural behavior simply because they know they are being observed by researchers.
- Demand Characteristics: Subtle cues in the experimental setup that reveal the researcher's hypothesis to participants, prompting them to alter their responses to fit (or thwart) expectations.
Types of Measurement & Experimental Validity
Beyond internal and external validity, researchers evaluate specific subtypes of measurement validity:
- Construct Validity: The degree to which an assessment tool or experimental protocol accurately measures the theoretical construct it claims to assess. Includes:
- Convergent Validity: The extent to which a test correlates strongly with other established tests measuring the same construct.
- Discriminant Validity: The extent to which a test does NOT correlate with tests measuring distinct, unrelated constructs.
- Criterion Validity: How well a score on a measurement tool predicts or correlates with a concrete real-world outcome. Includes:
- Predictive Validity: The ability of a test score to predict future performance (e.g., MCAT scores predicting first-year medical school GPA).
- Concurrent Validity: The agreement between a test score and a benchmark outcome measured at the exact same time.
- Content Validity: The extent to which a test comprehensively covers all dimensions and facets of the theoretical domain being evaluated.
- Face Validity: The superficial appearance of whether a test measures what it intends to measure (the weakest form of validity, assessed informally by non-experts).
Reliability: Consistency & Precision
Reliability refers to the consistency, stability, and reproducibility of a measurement tool across repeated administrations or across different observers.
Major Types of Reliability
- Test-Retest Reliability: The consistency of scores when the exact same measurement tool is administered to the same individuals at two different points in time (evaluating temporal stability).
- Inter-Rater Reliability: The degree of agreement or consistency between two or more independent raters/observers evaluating the same behavior or data set (quantified using metrics like Cohen's Kappa).
- Internal Consistency: The degree to which different items on the same test or survey produce similar results measuring the same underlying construct (commonly quantified using Cronbach's Alpha, (\alpha \ge 0.70) indicating acceptable internal reliability).
The Golden Rule of Reliability and Validity
Reliability is a NECESSARY, but NOT SUFFICIENT, condition for Validity.
A measurement instrument can be perfectly reliable (producing identical results every time) while being completely invalid (measuring the wrong variable). However, an instrument that is unreliable (producing random, inconsistent readings) can NEVER be valid.
Confounding, Mediating & Moderating Variables
In complex psychological and sociological models, researchers must categorize variables based on their structural role in the causal chain:
Definitions & Structural Roles
- Confounding Variable (Confounder): An unmeasured extraneous variable that independently influences BOTH the independent variable and the dependent variable. Confounders create a spurious association or obscure a true association.
- Mitigation: Randomization, matching, restriction, or statistical stratification/multivariate regression analysis.
- Mediating Variable (Mediator): A variable that sits directly on the causal pathway between the IV and DV. It explains HOW or WHY the independent variable produces the dependent variable (IV -> Mediator -> DV). If the mediator is removed or controlled, the observed relationship between IV and DV diminishes or disappears.
- Moderating Variable (Moderator): A variable that modifies the strength or direction of the relationship between the IV and DV. It answers WHEN or FOR WHOM the relationship holds true (e.g., an intervention works effectively in adults but fails in pediatric populations—age acts as a moderator).
Comparative Matrix of Variable Types
| Variable Type | Position in Causal Chain | Answers Question | Effect of Statistical Control | Example Scenario |
|---|---|---|---|---|
| Independent (IV) | Primary Cause / Input | What is being manipulated? | N/A (Main predictor variable) | Dosage of cognitive behavioral therapy (CBT). |
| Dependent (DV) | Primary Outcome | What is being measured? | N/A (Main outcome variable) | Reduction in clinical anxiety scores. |
| Confounding | Parallel influence on both IV and DV | Is the association real or spurious? | Eliminates spurious correlation | Socioeconomic status influencing both therapy access and anxiety levels. |
| Mediating | Intermediate step (IV -> Med -> DV) | HOW or WHY does IV affect DV? | Reduces or eliminates direct IV-DV relationship | Enhanced emotional regulation skills mediating CBT's effect on anxiety. |
| Moderating | Intersects and alters IV-DV link strength | WHEN or FOR WHOM does IV affect DV? | Reveals subgroup interaction effects | Biological sex modifying CBT efficacy (e.g., stronger effect in females). |
Systematic Biases in Psychological & Sociological Research
- Selection Bias: Non-random sampling causing the study population to differ systematically from the target population.
- Social Desirability Bias: Tendency of survey respondents to answer questions in a manner that will be viewed favorably by others, underreporting stigmatized behaviors.
- Acquiescence Bias: Tendency of respondents to agree with all statements ("yea-saying") regardless of content.
- Experimenter Expectancy Bias (Rosenthal Effect): Unconscious researcher body language or tone communicating expected outcomes to subjects, influencing their behavior to conform to hypotheses.
- Pygmalion Effect: Phenomenon where higher expectations placed on individuals (e.g., students by teachers) lead to an increase in performance.
- Impression Management: Active conscious strategies individuals employ to control how they are perceived by researchers or peers.
Worked MCAT Application Scenario
Scenario: A trial tests an online mindfulness app on workplace burnout among corporate employees. Employees who completed the 8-week app module showed significant reductions in burnout symptoms. However, critics point out three issues: (1) Employees self-selected into the app module; (2) The app reduced burnout primarily by increasing sleep quality; and (3) The app reduced burnout significantly in young employees, but had no effect on senior executives.
MCAT Variable Classification:
- Self-selection into the module introduces Selection Bias (a major threat to internal validity). Sicker or less motivated employees may have opted out.
- Sleep Quality acts as a Mediating Variable (App -> Sleep Quality -> Burnout Reduction), explaining how the app works.
- Employee Age / Job Rank acts as a Moderating Variable, altering the strength of the app's effect across subgroups.
A standardized personality inventory administered to 500 medical students yields nearly identical trait scores when retaken 6 months later. However, researchers discover that the inventory fails to predict actual clinical performance or bedside manner. This assessment tool demonstrates which of the following?
A study finds that higher levels of chronic work stress correlate with increased rates of coronary artery disease. Further statistical analysis reveals that chronic work stress elevates systemic inflammation (measured by C-reactive protein), which directly causes vascular endothelial damage leading to disease onset. When systemic inflammation is controlled for, the direct relationship between stress and coronary disease disappears. In this study, systemic inflammation acts as which type of variable?
Industrial psychologists evaluate worker productivity in a manufacturing plant following the installation of new ambient lighting. Productivity increases sharply during the 4-week observation period. However, subsequent analysis reveals that productivity increased equally in control rooms where lighting was left unchanged, because workers in both rooms knew they were participating in an active research study. This phenomenon is an example of which of the following?
Researchers investigate the effect of a novel anti-hypertensive medication on blood pressure. They find that the medication significantly reduces systolic blood pressure in elderly patients (ages 65+), but demonstrates no statistically significant effect in young adults (ages 18-35). In this experimental framework, patient age functions as which of the following?