20.2 Psychological Research Methods

Key Takeaways

  • Experimental designs manipulate an independent variable while controlling confounds; correlational designs measure relationships without establishing causation.
  • Internal validity concerns causal claims within the study; external validity concerns generalizability; reliability is reproducibility of measurement.
  • Probability sampling (simple random, stratified) supports generalization; non-probability sampling (convenience, snowball) is common but limits external validity.
  • The Belmont Report and IRB oversight require respect for persons, beneficence, and justice; informed consent and confidentiality are mandatory except in specific exempt categories.
  • Common threats to validity include selection bias, history, maturation, testing effects, and experimenter expectancy (addressed by double-blind designs).
Last updated: August 2026

Psychological Research Methods on the PA-CAT

Quick Answer: The PA-CAT Bulletin of Information, rev. 20240815 expects you to distinguish experimental from correlational designs, identify independent/dependent variables, evaluate validity and reliability, recognize sampling methods, and apply research ethics (IRB, informed consent). Expect items that give a study scenario and ask what conclusion is warranted.

Experimental vs Correlational Designs

FeatureExperimentalCorrelational
ManipulationYes — IV is controlled by researcherNo — variables are measured
Random assignmentYesNo
Causal inferencePermittedNot permitted
Typical question"Does X cause Y?""Are X and Y related, and how strongly?"

In a true experiment, the researcher manipulates the independent variable (IV) and measures the dependent variable (DV) while controlling confounds. A correlational study quantifies the relationship between two measured variables using Pearson r (−1 to +1) but cannot establish direction or causation.

Variables and Hypotheses

A hypothesis is a testable prediction. The null hypothesis (H₀) states no effect; the alternative hypothesis (H₁) states an effect exists. Type I error (α) is rejecting a true null (false positive); Type II error (β) is failing to reject a false null (false negative). Statistical power (1 − β) increases with sample size and effect size.

Sampling Methods

Sampling TypeMethodStrength
Simple randomEvery member has equal chanceMinimizes selection bias
Stratified randomRandom within subgroupsEnsures subgroup representation
ClusterRandom groups, then all membersPractical for large populations
ConvenienceWhoever is availableLow cost; weak generalizability
SnowballReferrals from participantsUseful for hidden populations

Probability sampling supports population inference; non-probability sampling limits external validity.

Validity and Reliability

Validity is whether a study measures what it claims; reliability is whether measurement is reproducible.

  • Internal validity — extent to which the IV, not confounds, caused the DV effect.
  • External validity — generalizability to other people, settings, times.
  • Construct validity — the operationalization captures the intended construct.
  • Criterion validity — correlates with an outcome measure (concurrent or predictive).
  • Content validity — items sample the full domain.
  • Inter-rater reliability — agreement across observers.
  • Test-retest reliability — stability over time.
  • Internal consistency — items on a scale measure the same thing (Cronbach's α).

Threats to Internal Validity

Campbell and Stanley's classic threats include selection (groups differ at baseline), history (external event during study), maturation (natural change over time), testing (practice effects), instrumentation (measurement drift), and experimenter expectancy (Rosenthal effect—expectations influence results). Double-blind designs, random assignment, and placebo controls mitigate these threats.

Research Designs

  • Cross-sectional — one time point; compares groups (e.g., by age).
  • Longitudinal — same participants over time; stronger causal inference.
  • Case-control — compares cases with a condition to controls without.
  • Quasi-experimental — manipulation without random assignment (common in education/clinical settings).
  • Single-subject (case study) — in-depth analysis of one individual.

Research Ethics

The Belmont Report (1979) established three principles:

  1. Respect for persons — autonomy; informed consent; protection of vulnerable groups.
  2. Beneficence — maximize benefit, minimize harm.
  3. Justice — fair distribution of research burdens and benefits.

An Institutional Review Board (IRB) reviews human-subjects research. Informed consent requires disclosure of purpose, procedures, risks, benefits, voluntariness, and the right to withdraw. Confidentiality and anonymity are distinct: confidentiality means the researcher can link data to identity but promises not to disclose; anonymity means no identifying link exists.

Deception is permitted only when methodologically necessary and must be followed by debriefing. Animal research is governed by IACUC oversight and the principle of the three Rs: replacement, reduction, refinement.

Validity and Reliability

Reliability is the consistency of measurement; validity is whether a measure actually captures the intended construct. A measure can be reliable but invalid (a scale that reads 3 kg heavy every time). The core distinctions a PA-CAT item tests:

ConceptDefinition
Internal validityDegree to which a design supports a causal claim within the study
External validityGeneralizability to other populations, settings, and times
Construct validityWhether the operationalization reflects the theoretical construct
Test-retest reliabilityConsistency of a measure across time
Inter-rater reliabilityConsistency across different observers

Threats to Validity

Confounding (a third variable explains the association), selection bias (groups differ at baseline), history and maturation (external events or natural change during the study), and experimenter expectancy (the Rosenthal effect) all threaten internal validity. Random assignment neutralizes selection and confounding; blinding limits expectancy effects. External validity is threatened by non-representative sampling (a college-student convenience sample) and artificial lab settings. The PA-CAT commonly presents a study scenario and asks which threat best explains a flaw or which design change would remove it, so practice mapping each threat to its specific fix.

Operationalization, Validity Threats, and Reliability

Operationalization is the step that turns an abstract construct into a measurable variable, and PA-CAT items often hinge on whether it was done well. A study on test anxiety is not researchable until the investigator specifies exactly how it is measured: a self-report Likert scale, heart rate during the exam, or salivary cortisol before testing. Each choice yields a different operational definition. The independent variable (IV) is the factor the researcher manipulates, the dependent variable (DV) is the measured outcome, and control variables are held constant to rule out alternative explanations. In a worked example, a researcher tests whether a 20-minute mindfulness exercise (IV: mindfulness vs quiet reading) reduces state anxiety (DV: STAI score). Holding room temperature, time of day, and tester constant removes confounds; random assignment to conditions neutralizes pre-existing group differences. Without random assignment the design collapses into quasi-experimental, and a causal claim is no longer warranted.

The PA-CAT consistently pairs each named threat with the design fix that removes it, so memorize the threat-fix pairings rather than threats alone. Selection bias (groups differ at baseline) is fixed by random assignment. History (an external event coincides with treatment) is fixed by a control group experiencing the same event. Maturation (natural change over time) is fixed by a comparison group. Testing (practice effects from repeated measurement) is fixed by alternate forms. Instrumentation (observer drift or recalibrated equipment) is fixed by standardized training and blinding. Attrition (differential dropout) is fixed by intent-to-treat analysis. Experimenter expectancy (Rosenthal effect) is fixed by double-blind procedures where neither participant nor observer knows condition assignment.

Reliability is consistency of measurement, in three flavors the exam distinguishes: test-retest reliability (stability across time), inter-rater reliability (agreement across observers), and internal consistency (Cronbach alpha, where alpha greater than 0.7 is acceptable). A scale can be reliably wrong: a biased bathroom scale reading 3 kg heavy every morning has high test-retest reliability but zero validity. Construct validity asks whether the operationalization captures the intended theoretical construct; internal validity asks whether the IV, not a confound, produced the DV change; external validity asks whether results generalize to other populations, settings, and times.

The canonical PA-CAT trap presents a correlational finding and a causal conclusion. A researcher reports that students who volunteer more have higher GPAs and claims volunteering raises grades. This is a correlational design: no variable was manipulated, no random assignment occurred, and a third variable (conscientiousness, socioeconomic status) or reverse causation could explain the association. Correlation does not imply causation because it cannot rule out confounds or establish direction. A true experiment adds two things: manipulation of the IV and random assignment of participants to conditions. Random assignment distributes individual differences across groups probabilistically, making groups equivalent at baseline and isolating the IV as the cause of any DV difference. When a stem asks what design upgrade would permit a causal claim, the answer is almost always randomly assign participants to conditions—not a larger sample, more measures, or a longitudinal follow-up, none of which substitute for random assignment.

Common Statistical Decision Thresholds (typical convention)
Test Your Knowledge

A researcher finds that students who sleep more score higher on the PA-CAT and concludes that more sleep causes higher scores. This conclusion is most vulnerable to which limitation?

A
B
C
D
Test Your Knowledge

A study uses a new depression scale that produces nearly identical scores on retest two weeks later but correlates poorly with clinician-rated depression. The scale has high:

A
B
C
D