14.2 Between-Subjects, Within-Subjects, Quasi-Experimental, and Correlational Designs
Key Takeaways
True experimental designs require the active manipulation of at least one independent variable and the random assignment of participants to conditions, establishing the gold standard for causal inference.
Between-subjects designs isolate conditions across independent participant cohorts to eliminate order and carryover effects, whereas within-subjects (repeated-measures) designs maximize statistical power by using each subject as their own control, necessitating counterbalancing.
Factorial designs cross two or more independent variables to evaluate both main effects and interaction effects; a significant interaction indicates that the effect of one independent variable varies depending on the level of another.
Quasi-experiments lack random assignment and rely on designs such as non-equivalent control groups, interrupted time-series, and regression-discontinuity to investigate phenomena where manipulation is impossible or unethical.
Developmental research balances cross-sectional designs (vulnerable to cohort effects), longitudinal designs (vulnerable to selective attrition and testing effects), and sequential designs (Schaie's design, which explicitly isolates age, cohort, and period effects).
Between-Subjects, Within-Subjects, Quasi-Experimental, and Correlational Designs
Designing an empirical study requires selecting an architectural framework that balances internal validity, statistical power, practical feasibility, and generalizability. Methodologists categorize research designs along a continuum of experimental control—from true experiments capable of establishing definitive causal links, to quasi-experimental approximations, to observational and correlational paradigms designed for ecological description and prediction. Mastery of these structural designs and their analytical signatures is essential for success on the GRE Psychology Subject Test.
1. True Experimental Designs: Between-Subjects vs. Within-Subjects
A true experiment is defined by two non-negotiable architectural criteria: (1) active, direct manipulation of one or more independent variables by the researcher, and (2) random assignment of participants to treatment conditions. Experiments generally configure participant exposure in one of two fundamental arrangements:
Between-Subjects (Independent Groups) Within-Subjects (Repeated Measures)
Participant Pool Participant Pool
┌─────────┴─────────┐ │
▼ ▼ ▼
[ Group A ] [ Group B ] [ All Participants ]
(Condition 1) (Condition 2) │
┌───────┴───────┐
▼ ▼
Condition 1 ────► Condition 2
(Counterbalanced Order)
Between-Subjects Designs (Independent Groups)
In a between-subjects design, each participant is randomly assigned to one and only one level of the independent variable. Comparisons are conducted across distinct, independent groups of individuals:
- Advantages: Completely eliminates order effects, carryover effects, and fatigue. Participants remain naive to other conditions and alternative manipulations, minimizing demand characteristics.
- Disadvantages: Susceptible to individual difference error variance. Because different humans inhabit each cell, baseline differences in cognitive capacity, personality, or biology contribute to between-group variance (), lowering statistical power. Requires substantially larger total sample sizes ().
- Matched-Groups Design: A hybrid between-subjects variant designed to reduce individual difference variance. Participants are pretested on a relevant matching variable known to correlate strongly with the DV (e.g., baseline IQ, working memory capacity). Participants are ranked, paired into homogeneous clusters based on their scores, and then members of each pair are randomly assigned to different treatment conditions. This equalizes the groups on that critical attribute while maintaining between-subjects isolation.
Within-Subjects Designs (Repeated Measures)
In a within-subjects design, every participant is exposed to all levels of the independent variable in a sequential order. Comparisons are drawn within the same individuals across time:
- Advantages: Outstanding statistical power. Because each participant serves as their own experimental control, inter-individual baseline variance is mathematically partialled out of the error term (), vastly increasing sensitivity to detect true treatment effects. Economical: requires far fewer participants than a comparable between-subjects design.
- Disadvantages: Highly vulnerable to order effects and carryover effects.
- Order Effects: Changes in performance resulting from the ordinal position of a condition in the sequence, regardless of which treatment is administered. Includes practice effects (improvements due to task familiarization and learning) and fatigue/boredom effects (deterioration in performance due to cognitive exhaustion).
- Carryover Effects: Severe methodological confounds where the lingering cognitive, neurochemical, or psychological effects of an earlier condition persist and contaminate performance in a subsequent condition (e.g., testing an active pharmaceutical agent with a long metabolic half-life before testing a placebo).
Counterbalancing Techniques
To neutralize order effects in repeated-measures paradigms, researchers implement counterbalancing—systematically varying the order in which conditions are presented across participants:
[ COUNTERBALANCING STRATEGIES ]
│
┌────────────────────────────┴────────────────────────────┐
▼ ▼
[ Complete Counterbalancing ] [ Incomplete Counterbalancing ]
(All k! possible sequences administered) ├── Latin Square (1x per position)
├── Balanced Latin Square (precedes/follows 1x)
└── Block Randomization (Multi-trial blocks)
- Complete Counterbalancing: Every possible mathematical permutation of condition sequences is administered. The total number of required sequences is (factorial of the number of conditions, ). For 2 conditions: orders (). For 3 conditions: orders (). For 4 conditions: orders. For 5 conditions: orders. When , complete counterbalancing rapidly becomes logistically intractable due to excessive sample size requirements.
- Incomplete Counterbalancing — Latin Square Design: A square matrix where each condition appears exactly once in each ordinal row (participant sequence) and exactly once in each ordinal column (temporal position). While controlling for the ordinal position of conditions, a standard Latin square does not control for sequence-specific carryover effects.
- Balanced Latin Square: An advanced incomplete counterbalancing architecture ensuring that: (a) each condition appears in each ordinal position exactly once, and (b) each condition immediately precedes and immediately follows every other condition an equal number of times. This controls for both general order effects and immediate linear carryover effects. For an even number of conditions (), a balanced Latin square can be generated using the algorithmic sequence for the first row:
For conditions (), Row 1 is . Subsequent rows are generated by adding 1 to each element (with cycling back to 1):
- Row 1:
- Row 2:
- Row 3:
- Row 4:
(Note: If is an odd number, two distinct Latin squares must be combined to achieve balance). 4. Block Randomization: Used in within-subjects designs where each condition is presented multiple times to each participant (e.g., psychophysics, reaction time tasks). Each 'block' contains one randomized instance of every condition, preventing participants from anticipating upcoming conditions.
2. Factorial Experimental Designs
Real-world psychological phenomena rarely stem from isolated independent variables acting in a vacuum. Factorial designs allow researchers to simultaneously manipulate two or more independent variables (termed factors) within a single integrated experiment.
Factorial Nomenclature and Architecture
Factorial designs are described using numerical notation that indicates both the number of independent variables and the number of levels per variable:
- A Factorial Design has two factors, each possessing two levels, yielding a total of distinct experimental conditions (cells).
- A Design has three factors; Factor A has 2 levels, Factor B has 3 levels, and Factor C has 4 levels, yielding cells.
Design Classifications by Subject Assignment
- Between-Subjects Factorial: Completely independent groups of participants populate every cell of the matrix.
- Within-Subjects Factorial: Every participant completes every combination of all factor levels.
- Mixed Factorial Design (Split-Plot Design): Contains at least one between-subjects factor and at least one within-subjects factor. For example, testing the cognitive performance of young vs. older adults (between-subjects subject variable) under three different memory load conditions (within-subjects repeated factor).
Main Effects and Interaction Effects
A factorial ANOVA evaluates two distinct types of statistical effects:
- Main Effect: The overall statistical effect of a single independent variable on the dependent variable, averaged across all levels of all other independent variables. In a design, there are two potential main effects: the main effect of Factor A and the main effect of Factor B.
- Interaction Effect: Occurs when the effect of one independent variable on the dependent variable changes as a function of the level of another independent variable. An interaction signifies moderation—the impact of IV-A is dependent upon IV-B.
No Interaction Interaction Present
(Parallel Slopes) (Non-Parallel / Crossing Slopes)
DV DV
▲ ▲
│ ─── Condition B1 │ ─── Condition B1
│ ─── Condition B2 │ ─── Condition B2
│ │ X (Crossed)
│ │ / \
└──────────────────► └───────────/───\──────►
A1 A2 A1 A2
(Factor A) (Factor A)
Graphical Diagnostics and Interpretive Rules
- Parallel Lines: When the line segments depicting cell means on a Cartesian interaction plot are parallel (or approximately parallel), no interaction exists; the effect of Factor A is identical across all levels of Factor B.
- Non-Parallel Lines: Diverging, converging, or intersecting lines indicate the presence of an interaction effect.
- Crossed (Disordinal) Interaction: The lines cross each other, indicating that the direction of the effect of Factor A completely reverses across levels of Factor B.
- Spreading (Ordinal) Interaction: The lines diverge from a common point, indicating that Factor A has an effect at one level of Factor B, but no effect (or an attenuated effect) at another level.
- Interpretive Rule of Interactions: A significant interaction qualifies, supersedes, and limits the interpretability of main effects. When a significant interaction is present, researchers cannot make broad, unqualified general statements about a main effect without examining the simple main effects at each specific level of the moderating factor.
3. Quasi-Experimental and Non-Experimental Methodologies
In many empirical contexts, random assignment is ethically indefensible (e.g., assigning pregnant mothers to consume alcohol), physically impossible (e.g., assigning biological age or clinical schizophrenia), or logistically unfeasible. In such cases, researchers deploy quasi-experimental and non-experimental architectures.
Quasi-Experimental Designs
Quasi-experiments resemble true experiments in that they manipulate an IV or study a treatment intervention, but they lack true random assignment of participants to conditions:
- Non-Equivalent Control Group Pretest-Posttest Design:
- Architecture: An intact, pre-existing group (e.g., Class A) receives an intervention (), while a separate intact group (e.g., Class B) serves as a comparison control ().
- Methodological Utility: The pretest () establishes whether the two intact groups were comparable on the DV prior to treatment.
- Vulnerability: Susceptible to selection interactions (e.g., selection-maturation: one group may mature or learn at a faster natural rate than the other, creating posttest differences independent of the treatment).
- Interrupted Time-Series Design:
- Architecture: A single group or population is measured repeatedly across extensive intervals both before and after the introduction of an intervention or environmental event:
- Methodological Utility: By establishing a clear baseline trajectory (), the researcher can detect and rule out pre-existing secular maturation trends, cyclical seasonal variations, and statistical regression. If an abrupt discontinuity or persistent slope shift appears precisely at , confidence in a causal treatment effect is substantially enhanced.
- Regression-Discontinuity Design:
- Architecture: Participants are assigned to experimental versus control conditions strictly based on whether their score on a quantitative pre-assignment metric falls above or below an exact mathematical cutoff threshold (e.g., students scoring below 70 on a reading test receive specialized tutoring; students scoring 70 or above receive standard instruction).
- Methodological Utility: If a sharp, discontinuous vertical jump in posttest performance occurs precisely at the cutoff value along the continuous regression line, it provides strong causal evidence of treatment efficacy, approaching the internal validity of a randomized trial.
Observational and Non-Experimental Research
- Naturalistic Observation: Passive, systematic observation and quantitative coding of behavior in natural ecological habitats without researcher intervention (e.g., Jane Goodall's chimpanzee ethograms). Maximizes ecological validity, but sacrifices experimental control; risks observer bias and participant reactivity.
- The Case Study Method (Idiographic Approach): An exhaustive, in-depth empirical investigation of a single individual, clinical patient, or unique event. Essential for examining rare neurological lesions or psychopathological anomalies that cannot be experimentally manufactured (e.g., Phineas Gage's ventromedial prefrontal lesion revealing executive emotional regulation; Patient H.M.'s bilateral medial temporal lobectomy revealing the memory consolidation role of the hippocampus). Limitations: zero population generalizability, lack of control comparisons, high risk of researcher confirmation bias.
- Archival Research: Analysis of pre-existing public records, historical documents, medical databases, or census registries. Completely non-reactive (participants cannot alter behavior because data was previously recorded), but constrained by missing records, non-standardized historical metrics, and correlational limitations.
- Survey Research: Quantifying self-reported attitudes, beliefs, and behaviors across sample cohorts. Vulnerable to non-response bias, sampling bias, framing effects, and response sets (acquiescence bias [the tendency to agree with all statements] and social desirability bias).
4. Developmental Research Designs
Developmental psychology investigates how cognitive, affective, and physiological processes change across the human lifespan. Studying chronological age presents unique methodological challenges because age is a subject variable that cannot be manipulated or randomly assigned. Methodologists utilize three primary architectures to study developmental trajectories:
Cross-Sectional: [ Cohort A (Age 20) ] ─┐
[ Cohort B (Age 40) ] ─┼──► All measured at single Time Point (Year 2026)
[ Cohort C (Age 60) ] ─┘ (Confounded by COHORT EFFECTS)
Longitudinal: [ Cohort A (Born 1980) ] ──► Tested 2000 (Age 20) ──► Tested 2020 (Age 40) ──► Tested 2040 (Age 60)
(Confounded by ATTRITION, PRACTICE, & PERIOD EFFECTS)
Sequential: Combines both: Tracks multiple cohorts longitudinally across multiple measurement waves.
| Research Design | Architectural Structure | Primary Advantages | Critical Methodological Confounders |
|---|---|---|---|
| Cross-Sectional Design | Compares multiple distinct age cohorts simultaneously at a single point in chronological time. | Highly efficient; rapid data collection; low financial cost; zero participant attrition. | Cohort Effects (Generational Confounding): Differences between age groups may reflect differences in historical, cultural, educational, or technological upbringing rather than true biological/cognitive developmental aging. |
| Longitudinal Design | Tracks the exact same cohort of participants repeatedly across an extended chronological timespan. | Directly observes intra-individual developmental change and stability over time; eliminates cohort differences. | Selective Attrition (Mortality): Non-random participant drop-out (e.g., healthier, wealthier individuals survive/remain, biasing older samples); Testing/Practice Effects from repeated test exposure; Period Effects (Historical Confounding). |
| Sequential / Cross-Sequential Design | Recruits multiple age cohorts simultaneously and follows all of them repeatedly across multiple measurement waves (K. Warner Schaie's Most Efficient Design). | Disentangles the three collinear developmental parameters: Chronological Age, Birth Cohort, and Time of Measurement (Period Effect). | Extremely expensive, logistically complex, requires decades of sustained funding and sophisticated hierarchical linear modeling. |
Note
The Classic Intelligence Debate (Schaie's Seattle Longitudinal Study): Early cross-sectional studies suggested that general fluid intelligence begins an irreversible, steep decline starting in a person's mid-20s. However, when K. Warner Schaie deployed sequential designs in the Seattle Longitudinal Study, he proved that this apparent early decline was an artifact of cohort effects—older cohorts had received fewer years of formal education and poorer childhood healthcare than younger cohorts. Longitudinal tracking revealed that cognitive capabilities remain remarkably stable until well into the 60s and 70s.
A psycholinguist wants to investigate the reaction time to four different font types (A, B, C, D) using a within-subjects repeated-measures design. To control for both the ordinal position of each condition and immediate linear carryover effects, the researcher constructs a Balanced Latin Square. Which of the following conditions must be satisfied by this design?
Each condition must appear only once across the entire experiment, distributed randomly among four independent participant cohorts.
Participants are pretested on reading speed and then permanently assigned to a single, non-repeating font condition based on their baseline score.
Every participant must be exposed to all twenty-four possible mathematical permutations of the four conditions.
Each condition appears in each ordinal position exactly once, and each condition precedes and follows every other condition an equal number of times.
A 2x2 factorial experiment investigates the effects of Sleep Deprivation (Rest vs. Deprived) and Task Complexity (Simple vs. Complex) on mathematical error rates. The results indicate that Sleep Deprivation significantly increases error rates on Complex tasks, but has no measurable impact on Simple tasks. On an interaction plot, what graphical pattern demonstrates this finding?
Non-parallel line segments that diverge or cross, indicating a significant interaction effect
A single curvilinear inverted U-shaped function reflecting the Yerkes-Dodson arousal law
Two perfectly horizontal, overlapping lines indicating the absence of main effects
Two parallel line segments with identical positive slopes indicating equivalent main effects
A municipality introduces a mandatory law requiring drivers to use hands-free cellular devices. A traffic researcher analyzes the monthly automobile collision rates across the city for forty-eight months prior to the law's enactment and forty-eight months following its enactment. What research architecture does this study exemplify?
An interrupted time-series quasi-experimental design
A randomized factorial clinical trial
A cross-sectional cohort developmental survey
A matched-groups between-subjects laboratory experiment
An educational psychologist conducts a cross-sectional study comparing the technological literacy of 20-year-olds, 40-year-olds, and 60-year-olds in 2026, finding that 20-year-olds score significantly higher than 60-year-olds. The researcher concludes that advancing biological age causes a deterioration in technological aptitude. Why is this conclusion fundamentally flawed?
The study violated ethical guidelines by failing to randomly assign chronological ages to the participants in each group.
Cross-sectional studies inherently suffer from high rates of selective participant attrition.
The observed differences are likely confounded by generational cohort effects rather than true developmental aging.
The researcher failed to counter-balance the conditions using a double-blind Latin square matrix.
Sections you finish are checked off in the contents.