20.2 Psychological Research Methods
Key Takeaways
- Experimental designs manipulate an independent variable while controlling confounds; correlational designs measure relationships without establishing causation.
- Internal validity concerns causal claims within the study; external validity concerns generalizability; reliability is reproducibility of measurement.
- Probability sampling (simple random, stratified) supports generalization; non-probability sampling (convenience, snowball) is common but limits external validity.
- The Belmont Report and IRB oversight require respect for persons, beneficence, and justice; informed consent and confidentiality are mandatory except in specific exempt categories.
- Common threats to validity include selection bias, history, maturation, testing effects, and experimenter expectancy (addressed by double-blind designs).
Psychological Research Methods on the PA-CAT
Quick Answer: The PA-CAT Bulletin of Information, rev. 20240815 expects you to distinguish experimental from correlational designs, identify independent/dependent variables, evaluate validity and reliability, recognize sampling methods, and apply research ethics (IRB, informed consent). Expect items that give a study scenario and ask what conclusion is warranted.
Experimental vs Correlational Designs
| Feature | Experimental | Correlational |
|---|---|---|
| Manipulation | Yes — IV is controlled by researcher | No — variables are measured |
| Random assignment | Yes | No |
| Causal inference | Permitted | Not permitted |
| Typical question | "Does X cause Y?" | "Are X and Y related, and how strongly?" |
In a true experiment, the researcher manipulates the independent variable (IV) and measures the dependent variable (DV) while controlling confounds. A correlational study quantifies the relationship between two measured variables using Pearson r (−1 to +1) but cannot establish direction or causation.
Variables and Hypotheses
A hypothesis is a testable prediction. The null hypothesis (H₀) states no effect; the alternative hypothesis (H₁) states an effect exists. Type I error (α) is rejecting a true null (false positive); Type II error (β) is failing to reject a false null (false negative). Statistical power (1 − β) increases with sample size and effect size.
Sampling Methods
| Sampling Type | Method | Strength |
|---|---|---|
| Simple random | Every member has equal chance | Minimizes selection bias |
| Stratified random | Random within subgroups | Ensures subgroup representation |
| Cluster | Random groups, then all members | Practical for large populations |
| Convenience | Whoever is available | Low cost; weak generalizability |
| Snowball | Referrals from participants | Useful for hidden populations |
Probability sampling supports population inference; non-probability sampling limits external validity.
Validity and Reliability
Validity is whether a study measures what it claims; reliability is whether measurement is reproducible.
- Internal validity — extent to which the IV, not confounds, caused the DV effect.
- External validity — generalizability to other people, settings, times.
- Construct validity — the operationalization captures the intended construct.
- Criterion validity — correlates with an outcome measure (concurrent or predictive).
- Content validity — items sample the full domain.
- Inter-rater reliability — agreement across observers.
- Test-retest reliability — stability over time.
- Internal consistency — items on a scale measure the same thing (Cronbach's α).
Threats to Internal Validity
Campbell and Stanley's classic threats include selection (groups differ at baseline), history (external event during study), maturation (natural change over time), testing (practice effects), instrumentation (measurement drift), and experimenter expectancy (Rosenthal effect—expectations influence results). Double-blind designs, random assignment, and placebo controls mitigate these threats.
Research Designs
- Cross-sectional — one time point; compares groups (e.g., by age).
- Longitudinal — same participants over time; stronger causal inference.
- Case-control — compares cases with a condition to controls without.
- Quasi-experimental — manipulation without random assignment (common in education/clinical settings).
- Single-subject (case study) — in-depth analysis of one individual.
Research Ethics
The Belmont Report (1979) established three principles:
- Respect for persons — autonomy; informed consent; protection of vulnerable groups.
- Beneficence — maximize benefit, minimize harm.
- Justice — fair distribution of research burdens and benefits.
An Institutional Review Board (IRB) reviews human-subjects research. Informed consent requires disclosure of purpose, procedures, risks, benefits, voluntariness, and the right to withdraw. Confidentiality and anonymity are distinct: confidentiality means the researcher can link data to identity but promises not to disclose; anonymity means no identifying link exists.
Deception is permitted only when methodologically necessary and must be followed by debriefing. Animal research is governed by IACUC oversight and the principle of the three Rs: replacement, reduction, refinement.
Validity and Reliability
Reliability is the consistency of measurement; validity is whether a measure actually captures the intended construct. A measure can be reliable but invalid (a scale that reads 3 kg heavy every time). The core distinctions a PA-CAT item tests:
| Concept | Definition |
|---|---|
| Internal validity | Degree to which a design supports a causal claim within the study |
| External validity | Generalizability to other populations, settings, and times |
| Construct validity | Whether the operationalization reflects the theoretical construct |
| Test-retest reliability | Consistency of a measure across time |
| Inter-rater reliability | Consistency across different observers |
Threats to Validity
Confounding (a third variable explains the association), selection bias (groups differ at baseline), history and maturation (external events or natural change during the study), and experimenter expectancy (the Rosenthal effect) all threaten internal validity. Random assignment neutralizes selection and confounding; blinding limits expectancy effects. External validity is threatened by non-representative sampling (a college-student convenience sample) and artificial lab settings. The PA-CAT commonly presents a study scenario and asks which threat best explains a flaw or which design change would remove it, so practice mapping each threat to its specific fix.
Operationalization, Validity Threats, and Reliability
Operationalization is the step that turns an abstract construct into a measurable variable, and PA-CAT items often hinge on whether it was done well. A study on test anxiety is not researchable until the investigator specifies exactly how it is measured: a self-report Likert scale, heart rate during the exam, or salivary cortisol before testing. Each choice yields a different operational definition. The independent variable (IV) is the factor the researcher manipulates, the dependent variable (DV) is the measured outcome, and control variables are held constant to rule out alternative explanations. In a worked example, a researcher tests whether a 20-minute mindfulness exercise (IV: mindfulness vs quiet reading) reduces state anxiety (DV: STAI score). Holding room temperature, time of day, and tester constant removes confounds; random assignment to conditions neutralizes pre-existing group differences. Without random assignment the design collapses into quasi-experimental, and a causal claim is no longer warranted.
The PA-CAT consistently pairs each named threat with the design fix that removes it, so memorize the threat-fix pairings rather than threats alone. Selection bias (groups differ at baseline) is fixed by random assignment. History (an external event coincides with treatment) is fixed by a control group experiencing the same event. Maturation (natural change over time) is fixed by a comparison group. Testing (practice effects from repeated measurement) is fixed by alternate forms. Instrumentation (observer drift or recalibrated equipment) is fixed by standardized training and blinding. Attrition (differential dropout) is fixed by intent-to-treat analysis. Experimenter expectancy (Rosenthal effect) is fixed by double-blind procedures where neither participant nor observer knows condition assignment.
Reliability is consistency of measurement, in three flavors the exam distinguishes: test-retest reliability (stability across time), inter-rater reliability (agreement across observers), and internal consistency (Cronbach alpha, where alpha greater than 0.7 is acceptable). A scale can be reliably wrong: a biased bathroom scale reading 3 kg heavy every morning has high test-retest reliability but zero validity. Construct validity asks whether the operationalization captures the intended theoretical construct; internal validity asks whether the IV, not a confound, produced the DV change; external validity asks whether results generalize to other populations, settings, and times.
The canonical PA-CAT trap presents a correlational finding and a causal conclusion. A researcher reports that students who volunteer more have higher GPAs and claims volunteering raises grades. This is a correlational design: no variable was manipulated, no random assignment occurred, and a third variable (conscientiousness, socioeconomic status) or reverse causation could explain the association. Correlation does not imply causation because it cannot rule out confounds or establish direction. A true experiment adds two things: manipulation of the IV and random assignment of participants to conditions. Random assignment distributes individual differences across groups probabilistically, making groups equivalent at baseline and isolating the IV as the cause of any DV difference. When a stem asks what design upgrade would permit a causal claim, the answer is almost always randomly assign participants to conditions—not a larger sample, more measures, or a longitudinal follow-up, none of which substitute for random assignment.
A researcher finds that students who sleep more score higher on the PA-CAT and concludes that more sleep causes higher scores. This conclusion is most vulnerable to which limitation?
A study uses a new depression scale that produces nearly identical scores on retest two weeks later but correlates poorly with clinician-rated depression. The scale has high: