14.1 The Scientific Method, Variables, and Operational Definitions
Key Takeaways
Scientific inquiry in psychology rests on empiricism, determinism, parsimony (Occam's razor / Morgan's canon), and Karl Popper's criterion of falsifiability, which asserts that theories advance through conjecture and empirical refutation rather than inductive verification.
The open science movement and preregistration emerged in response to the replication crisis to systematically curb questionable research practices (QRPs) such as p-hacking, HARKing (Hypothesizing After Results are Known), and the file-drawer effect.
Psychological science pursues four progressive empirical objectives: description of baseline phenomena, prediction of lawful covariation, causal explanation through rigorous control, and practical behavioral application.
Operational definitions translate latent, unobservable theoretical constructs into concrete, standardized, and quantifiable operations, while dependent measures must be calibrated to avoid insensitive ceiling and floor effects.
Internal validity requires eliminating confounding variables—extraneous factors that systematically covary with levels of the independent variable—through experimental control techniques including elimination, constancy of conditions, balancing, and random assignment.
The Scientific Method, Variables, and Operational Definitions
Psychology transitioned from speculative philosophy into an independent scientific discipline through the rigorous application of empirical methodology. Scientific inquiry does not rely on subjective intuition, folk belief, or unconstrained rationalism; rather, it demands systematic observation, reproducible quantification, and formalized attempts to falsify theoretical assertions. For the GRE Subject Test in Psychology, students must command the epistemological foundations of science, understand how theoretical constructs are operationalized into quantifiable variables, and master the methodological controls that prevent extraneous factors from undermining experimental conclusions.
1. Epistemological Foundations of Psychological Science
Scientific psychology operates on several foundational epistemological and metaphysical assumptions that distinguish it from non-scientific systems of knowledge:
- Empiricism: The principle that valid knowledge about psychological phenomena must be derived from structured, objective, and publicly verifiable sensory observation and experimentation, rather than ungrounded deduction or appeal to authority.
- Determinism: The foundational metaphysical doctrine asserting that the universe is lawful and orderly, such that cognitive events, physiological processes, and behavioral responses are causally determined by antecedent physical and psychological factors that can be discovered through empirical research.
- Parsimony (Occam's Razor / Morgan's Canon): When two or more competing hypotheses account for an empirical phenomenon with equal explanatory adequacy, scientists must adopt the simpler explanation—that is, the one requiring the fewest unobserved assumptions and mechanisms. In comparative psychology and animal behavior, this principle is formalized as C. Lloyd Morgan's Canon (1894): 'In no case may we interpret an action as the outcome of the exercise of a higher psychical faculty, if it can be interpreted as the outcome of the exercise of one which stands lower in the psychological scale.'
- Falsifiability and Popperian Demarcation: In The Logic of Scientific Discovery (1934/1959), philosopher Karl Popper introduced falsifiability as the definitive criterion of demarcation separating genuine empirical science from pseudoscience and metaphysics. Popper argued that inductive verification is logically asymmetrical: no finite number of confirming empirical observations can definitively establish a universal affirmative hypothesis (e.g., observing one million white swans cannot conclusively prove the universal proposition that 'all swans are white'), whereas a single rigorous, contradictory empirical observation immediately refutes it (discovering a single black swan refutes the proposition). Consequently, scientific theories must generate precise, risky empirical predictions that explicitly specify what observational patterns would prove the theory wrong. In Popper's framework, science advances not through verification, but through conjecture and refutation; theories are never permanently proven—they are merely corroborated by repeatedly withstanding rigorous empirical attempts to falsify them.
2. The Replication Crisis and Open Science Reforms
Over the past two decades, psychological science experienced profound self-examination sparked by systemic failures to replicate landmark experimental findings. In a seminal 2015 large-scale replication initiative conducted by the Open Science Collaboration (led by Brian Nosek), 270 researchers attempted to replicate 100 high-profile psychological studies published in premier journals. While 97% of the original studies reported statistically significant results, only 36% to 39% of the direct replications yielded statistically significant findings, accompanied by a mean reduction in effect size of roughly 50%.
[ THE REPLICATION CRISIS ]
│
┌───────────────────────┴───────────────────────┐
▼ ▼
[ Questionable Research Practices ] [ Methodological Remedies ]
├── p-Hacking / Data Dredging ├── Study Preregistration
├── HARKing (Post-hoc Hypotheses) ├── Registered Reports
├── Selective Reporting / Publication Bias ├── Open Data, Code, & Materials
└── Underpowered Sample Sizes └── Large-Scale Multi-Site Consortia
Questionable Research Practices (QRPs)
Methodologists identified several widespread research practices that artificially inflate Type I error rates (false positives):
- p-Hacking (Data Dredging): Exploiting 'researcher degrees of freedom' by iteratively testing multiple post-hoc covariates, selectively dropping outliers, testing different dependent measures, or continuously checking significance while adding participants until the p-value dips just below the conventional threshold of .
- HARKing (Hypothesizing After Results are Known): Termed by Norbert Kerr (1998), HARKing occurs when researchers examine exploratory post-hoc data trends and retroactively present them in the research report as if they were a priori, theoretically grounded hypotheses, misrepresenting inductive exploration as deductive confirmation.
- The File-Drawer Problem (Publication Bias): Documented by Robert Rosenthal (1979), this refers to the systematic tendency for academic journals to publish statistically significant findings () while non-significant, null results remain unpublished in investigators' file drawers. This skews the published empirical literature, distorting meta-analytic effect sizes.
Open Science Methodological Standards
In response, the international scientific community established rigorous open science conventions:
- Preregistration: Researchers write and publicly deposit a detailed research plan—specifying directional hypotheses, sample size calculations, stopping rules, variable operationalizations, and precise statistical models—in a permanent, time-stamped registry (such as the Open Science Framework [OSF]) prior to collecting any empirical data. Preregistration draws a strict, transparent line between confirmatory hypothesis testing and exploratory post-hoc analysis.
- Registered Reports: A publishing model where peer review occurs before data collection. If the theoretical rationale, experimental design, and statistical power calculations pass peer review, the journal grants in-principle acceptance, guaranteeing publication regardless of whether the eventual findings are statistically significant or null.
- Open Materials and Open Data: Publicly archiving raw data sets, stimulus scripts, analysis scripts (R, Python, syntax files), and laboratory protocols to facilitate direct audit, verification, and exact replication.
Note
On the GRE Subject Test in Psychology, questions frequently target the epistemological distinction between deductive falsification and inductive confirmation. Remember: A hypothesis can never be mathematically 'proven' true in empirical science because an alternative, yet-untested mechanism could always account for the observed covariation. Hypotheses are either rejected (falsified) or temporarily supported (corroborated).
3. The Four Primary Goals of Psychological Research
Empirical psychological investigations pursue four progressive, hierarchically organized scientific objectives:
[ 1. Description ] ──► Identifies 'What' occurs (Frequencies, base rates, phenotypes)
│
▼
[ 2. Prediction ] ──► Forecasts 'When' it occurs (Lawful covariations, correlations)
│
▼
[ 3. Explanation ] ──► Uncovers 'Why' it occurs (Causality via experimental isolation)
│
▼
[ 4. Application ] ──► Translates 'How' to intervene (Therapies, policy, optimization)
| Research Objective | Core Scientific Focus | Primary Research Methodologies | Representative Exemplar |
|---|---|---|---|
| 1. Description | Systematically defining, categorizing, and cataloging behavioral and psychological phenomena, their prevalence, and their qualitative dimensions. | Naturalistic observation, ethograms, descriptive surveys, structured clinical interviews. | Cataloging the precise behavioral stages of rapid eye movement (REM) sleep or detailing the symptom profile of a newly recognized clinical syndrome. |
| 2. Prediction | Identifying reliable, systematic statistical covariations between variables to forecast future behavior without necessarily identifying causal mechanisms. | Correlational designs, longitudinal tracking, cross-sectional surveys, predictive regression modeling. | Using high school grade point average and standardized aptitude test scores to forecast first-year undergraduate academic performance (). |
| 3. Explanation (Causation) | Identifying the definitive cause-and-effect mechanisms that produce a phenomenon, answering why changes in one variable directly trigger changes in another. | True randomized laboratory and field experiments manipulating independent variables with tight control. | Demonstrating that acute sleep deprivation causally impairs hippocampal long-term potentiation and working memory consolidation by manipulating sleep intervals. |
| 4. Application (Control) | Translating empirical causal discoveries into real-world interventions, behavioral modifications, clinical treatments, ergonomics, and educational public policies. | Applied behavioral analysis, randomized controlled clinical trials (RCTs), human factors engineering. | Designing cognitive-behavioral exposure protocols to treat phobias or redesigning cockpit control panels to prevent pilot attentional error. |
John Stuart Mill's Criteria for Establishing Causality
To progress from mere prediction to definitive causal explanation, researchers must satisfy three strict empirical conditions derived from philosopher John Stuart Mill's canons:
- Covariation of Cause and Effect: The putative cause and the observed outcome must be statistically correlated; when the cause is present or changed, the outcome must change accordingly.
- Temporal Precedence: The putative cause must precede the observed effect chronologically in time (). Correlational and cross-sectional designs typically fail this condition due to the directionality problem.
- Elimination of Plausible Alternative Explanations (Non-Spuriousness): The researcher must demonstrate that no extraneous variable () is the true underlying source of the observed relationship. True randomized experiments accomplish this through rigorous experimental control.
4. Taxonomy of Variables and Operationalization
In psychological research, a variable is any observable property, attribute, or condition that can take on different numerical values or categorical levels across individuals, environments, or time.
Operational Definitions: Bridging Constructs to Measurements
Psychological theories deal largely with latent hypothetical constructs—unobservable conceptual entities such as 'intelligence,' 'anxiety,' 'ego depletion,' 'selective attention,' or 'working memory capacity.' Because constructs cannot be measured directly, researchers must formulate operational definitions.
- Operational Definition: A precise, unambiguous specification of the concrete physical procedures, operations, or measurement tools utilized to produce, manipulate, or quantify a theoretical construct in a specific study. First articulated in physics by Percy Bridgman (1927), operationalism mandates that a scientific concept is synonymous with the corresponding set of empirical operations by which it is measured.
- Example: The latent construct 'acute anxiety' can be operationally defined as: (a) a self-report score exceeding 60 on the State-Trait Anxiety Inventory (STAI), (b) an elevation in salivary cortisol concentration , or (c) an increase in mean skin conductance response of over baseline following a stressor.
[ Latent Theoretical Construct ] ──► 'Working Memory Capacity'
│
▼ (Operationalization Process)
[ Concrete Empirical Metric ] ──► Backward Digit Span Score (Maximum digits recalled)
Independent Variables (IVs)
The independent variable (IV) is the factor that is hypothesized to exert a causal influence on the dependent variable. In an experimental design, the IV must have at least two discrete levels or conditions (e.g., experimental vs. control; 0 mg, 10 mg, 20 mg drug dosage):
- True Manipulated IV: An environmental, task, or instructional condition that the experimenter directly creates, alters, and assigns to participants (e.g., sleep duration, ambient room temperature, reward magnitude). True manipulation enables causal inference.
- Subject Variable (Attribute / Organismic Variable / Quasi-IV): An intrinsic, pre-existing biological, demographic, or psychological characteristic possessed by the participant prior to the study (e.g., biological sex, chronological age, DSM-5 clinical diagnosis, introversion vs. extraversion). Subject variables cannot be randomly assigned. When researchers compare groups based on subject variables, the design is quasi-experimental, and direct causal claims are logically prohibited because unmeasured confounding factors may covary with the organismic trait.
Dependent Variables (DVs) and Modalities
The dependent variable (DV) is the measurable behavioral, cognitive, or physiological outcome that is hypothesized to depend upon, or vary as a function of, the independent variable. DVs generally fall into three empirical modalities:
- Behavioral Measures: Direct observations of physical action, response latency (reaction time in milliseconds), accuracy/error rates, frequency of target behaviors, or endurance.
- Physiological Measures: Biological indices including autonomic responses (galvanic skin response/GSR, heart rate variability), electrophysiological activity (EEG event-related potentials [ERPs]), neuroendocrine markers (cortisol, alpha-amylase), and functional neuroimaging (fMRI BOLD response).
- Self-Report Measures: Questionnaires, Likert scales, semantic differentials, and structured interviews reflecting subjective introspections, affective states, or personal beliefs.
Measurement Sensitivity: Ceiling and Floor Effects
A critical methodological flaw in psychological measurement is inadequate sensitivity of the dependent measure, which can completely mask true differences between treatment conditions:
Ceiling Effect (Task Too Easy) Floor Effect (Task Too Difficult)
Score Score
100 ───[■]───[■]───[■] (Max Limit) 100 ───
(Condition A, B, C cluster)
0 ─── 0 ───[■]───[■]───[■] (Min Limit)
(Condition A, B, C cluster)
- Ceiling Effect: Occurs when an experimental task is so excessively easy, or the measurement scale threshold is set so low, that virtually all participants achieve the maximum possible score regardless of the experimental manipulation. For example, if a cognitive neuroscientist tests the memory effects of a nootropic drug using a task requiring participants to remember three simple common words, both the drug and placebo groups will score 100% correct, obscuring any genuine pharmacological enhancement.
- Floor Effect: Occurs when an experimental task is so profoundly difficult, or the measurement threshold is set so unrealistically high, that virtually all participants score at or near the absolute minimum score regardless of condition. For example, if an educational psychologist tests a novel math curriculum on second graders using college-level calculus problems, all students will score 0%, masking any instructional benefit.
5. Extraneous vs. Confounding Variables and Control Techniques
The fundamental challenge of experimental methodology is isolating the causal impact of the independent variable from all other extraneous influences:
- Extraneous Variable: Any variable in the research setting other than the specified independent variable that could conceivably influence the dependent variable. Extraneous variables that are unsystematic introduce random noise (unexplained error variance, ), reducing statistical power but not necessarily invalidating causal conclusions.
- Confounding Variable: An extraneous variable that covaries systematically with the levels of the independent variable. A confound provides a viable, plausible alternative explanation for the observed changes in the dependent variable, completely destroying the internal validity of the experiment. If Group A receives a drug in a quiet morning room while Group B receives a placebo in a noisy afternoon room, ambient noise and time of day are fatal confounds.
[ Independent Variable (IV) ] ──────► [ Dependent Variable (DV) ]
▲ ▲
│ │
└─────[ Confounding Variable ]───┘
(Systematically covaries with IV
and directly influences DV)
Experimental Control Techniques
Methodologists deploy four primary techniques to eliminate or neutralize extraneous and confounding variables:
| Control Method | Implementation Mechanism | Type of Variance Controlled | Practical Limitations |
|---|---|---|---|
| Elimination | Completely removing the extraneous variable from the experimental environment (e.g., conducting cognitive testing in a sound-attenuated, light-controlled testing booth). | Eliminates both systematic confounding and unsystematic random error variance. | Severely restricts ecological validity; many variables (e.g., participant anxiety, ambient weather) cannot be physically eliminated. |
| Constancy of Conditions | Holding the extraneous factor identical and invariant across all experimental conditions (e.g., testing every participant at exactly 09:00 AM, using the exact same scripted instructions, room temperature of 21°C). | Converts a potential confound into a constant, preventing it from covarying with the IV. | Limits external generalizability to that specific, narrow set of environmental parameters. |
| Balancing (Matching) | Ensuring that the distribution of an extraneous variable is mathematically equivalent across all treatment groups (e.g., ensuring exactly 50% males and 50% females in both experimental and control groups). | Distributes known extraneous variables equally, preventing systematic covariation. | Only controls for the specific extraneous variables that the experimenter explicitly identifies and measures in advance. |
| Random Assignment | Utilizing an objective random procedure (e.g., true random number generators) such that every participant has an equal, non-zero, and independent probability of being assigned to any treatment condition. | The foundational pillar of experimental design. Controls for both known and unknown participant variables by dispersing them randomly across conditions. | Requires an adequate sample size ( per cell) to achieve probabilistic equivalence; small samples risk chance baseline non-equivalence. |
Note
Random Selection vs. Random Assignment: This distinction is one of the most frequently tested concepts on the GRE Psychology exam:
- Random Selection (Sampling): Every member of a target population has an equal chance of being selected to participate in the study. Governs External Validity (population generalizability).
- Random Assignment (Allocation): Every recruited participant has an equal chance of being placed into a particular experimental condition. Governs Internal Validity (unambiguous causal inference).
According to Karl Popper's criterion of demarcation, which of the following statements best characterizes why a psychological theory qualifies as genuinely scientific?
It articulates clear, risky empirical predictions that can be subjected to tests capable of refuting the theory.
It relies strictly on mathematical equations to model human cognitive and neurobiological architecture.
It can incorporate any post-hoc observational finding by adjusting its underlying theoretical assumptions.
It has been verified by an overwhelming preponderance of empirical observations across diverse populations.
A cognitive psychologist designs a novel computerized test of spatial working memory to evaluate the efficacy of a memory-training program. After four weeks of training, both the experimental group and the active control group achieve mean scores of 99.2% and 98.9% correct, respectively. The researcher concludes that the training program has zero efficacy. What methodological vulnerability invalidates this conclusion?
The dependent measure demonstrated low construct validity due to a lack of operational definitions.
The experimenter introduced an extraneous subject variable that created a non-spurious correlation.
The study suffered from regression toward the mean among low-performing baseline participants.
The test suffered from an acute ceiling effect that masked genuine differences in working memory capacity.
In an experiment investigating the effect of caffeine on sustained attention, participants in the high-caffeine condition are tested at 8:00 AM by a male experimenter, while participants in the decaffeinated control condition are tested at 4:00 PM by a female experimenter. In this design, what term specifically describes the time of day and experimenter gender?
Random error variance that decreases statistical power without biasing treatment effects
True independent variables manipulated in a factorial matrix
Organismic subject variables that restrict external validity
Confounding variables that compromise internal validity
A research team wants to ensure that baseline individual differences in working memory capacity do not compromise causal inference in a between-subjects drug trial. Which of the following procedural controls provides the most robust protection against both identified and unidentified individual differences?
Pre-screening participants and eliminating anyone whose baseline working memory score deviates from the median
Using true random assignment to allocate recruited participants across experimental and control conditions
Selecting a convenience sample of university undergraduates who share identical educational backgrounds
Holding environmental conditions constant by testing all subjects in the same laboratory chamber
Sections you finish are checked off in the contents.