12.3 Research Methods in Social Science
Key Takeaways
- Independent variables are manipulated (or designated predictors); dependent variables are measured outcomes—confounds threaten causal claims.
- Experiments support cause–effect when random assignment and control are strong; correlational designs show association only (correlation ≠ causation).
- Reliability is consistency of measurement; validity is whether the measure captures the intended construct and supports intended inferences.
- Ethics require informed consent, risk minimization, and careful limits on deception; NMAT tests research literacy, not advanced statistics software skills.
12.3 Research Methods in Social Science
Quick Answer: Know what was manipulated (IV), what was measured (DV), whether the design was experimental or correlational, whether the sample supports generalization, whether measures are reliable and valid, and whether ethics were respected. Correlation never proves causation.
CEM’s Social Science block rewards scientific thinking about people—the same literacy that protects you from weak claims in media, advertising, and even early research readings in medical school. You do not need advanced regression modeling; you need clean concepts.
Variables and Hypotheses
A hypothesis is a testable prediction. Variables are measurable attributes:
| Term | Definition | Example |
|---|---|---|
| Independent variable (IV) | Factor manipulated or treated as predictor | Study method (flashcards vs. rereading) |
| Dependent variable (DV) | Outcome measured | Quiz score |
| Confounding variable | Uncontrolled factor that covaries with IV and could explain DV | Prior GPA differing by condition |
| Operational definition | Concrete procedure for measuring a construct | “Anxiety” = score on a named scale |
Control variables are held constant or statistically adjusted so they do not obscure the IV–DV link. Random assignment (experiments) balances unknown confounds across groups on average; random sampling (from a population) supports external validity (generalization). Do not confuse the two.
Experimental vs. Correlational Designs
Experiments
In a true experiment, the researcher manipulates at least one IV, uses control conditions, and ideally randomly assigns participants. When these conditions hold, causal inference is strongest: changes in the IV cause changes in the DV (within design limits).
- Between-subjects: different people in each condition.
- Within-subjects: same people experience multiple conditions (order effects need counterbalancing).
- Placebo / single- and double-blind procedures reduce expectancy and observer bias when applicable.
Correlational and descriptive studies
Correlational research measures naturally occurring variables without manipulation. A correlation coefficient (r) indexes direction and strength of linear association (−1 to +1). Positive correlation: both rise together; negative: one rises as the other falls; near zero: little linear association.
Correlation ≠ causation. Three classic threats:
- Directionality: Does A cause B or B cause A?
- Third variable: C causes both A and B.
- Spurious / chance associations in noisy data.
Descriptive methods (case study, naturalistic observation, surveys) map phenomena and generate hypotheses but rarely alone establish cause.
| Design | Manipulation? | Best claim strength |
|---|---|---|
| True experiment | Yes + random assignment | Causal (with caveats) |
| Quasi-experiment | IV-like groups without full random assignment | Weaker causal |
| Correlational | No | Association only |
| Descriptive | No | Description / exploration |
Sampling
A population is the full set of interest; a sample is the studied subset.
- Probability sampling (simple random, stratified, systematic) gives each unit a known chance of selection—supports statistical generalization when executed well.
- Nonprobability sampling (convenience, snowball, voluntary response) is common and useful for early work but limits external validity.
- Sample size affects precision and statistical power; larger is not automatically “more valid” if the sample is systematically biased.
- Volunteer bias and undercoverage of hard-to-reach groups matter in health surveys.
Reliability and Validity
| Concept | Question answered | Examples / subtypes |
|---|---|---|
| Reliability | Is the measure consistent? | Test–retest, inter-rater, internal consistency |
| Validity | Does it measure what it claims and support intended uses? | Content, criterion, construct; internal vs. external validity of studies |
A scale can be reliable but not valid (consistently wrong construct). A study can have strong internal validity (tight causal control) but weak external validity (artificial lab). NMAT stems often ask which threat applies (confound, nonrandom sample, demand characteristics).
Demand characteristics are cues that lead participants to guess the hypothesis and alter behavior. Social desirability bias skews self-reports toward looking good. Observer bias distorts coding when raters know expected outcomes—hence blind coding and structured instruments.
Ethics in Human Research
Core principles (aligned with modern institutional review norms):
- Respect for persons: voluntary participation; informed consent with understandable information about procedures, risks, benefits, and withdrawal rights.
- Beneficence: minimize harm; risk–benefit balance; protect confidentiality/privacy of data.
- Justice: fair participant selection; avoid exploiting vulnerable groups without justification and safeguards.
Deception is sometimes used when full foreknowledge would destroy the study, but it is limited: scientific necessity, no unreasonable harm, and typically debriefing afterward. Historical abuses (and even well-known classic social psychology studies) drove today’s stricter oversight. NMAT may ask what ethical step was missing (no consent, no debrief, coercion).
Descriptive Statistics (Conceptual)
| Measure | Meaning | When it shines / fails |
|---|---|---|
| Mean | Arithmetic average | Sensitive to extreme outliers |
| Median | Middle value | Robust to outliers; skewed income-like data |
| Mode | Most frequent value | Categories; multimodal distributions |
| Range / variance / SD | Spread | High SD → more variability around mean |
Know that a skewed distribution pulls the mean toward the tail more than the median. Bar graphs vs. histograms, and reading axes, are quantitative-transfer skills; Social Science items stay conceptual.
Statistical significance (e.g., p < .05 culture) means results are unlikely under a null model—not that the effect is large or important. Effect size and practical significance are separate ideas; if an option confuses “significant” with “huge real-world impact,” reject it.
How NMAT-Style Items Test Research Literacy
Expect stems such as:
- “Which variable is independent?”
- “Why can the researcher not conclude causation?”
- “Which sampling problem is illustrated?”
- “Which ethical principle was violated?”
- “Which measure of central tendency is least affected by an extreme score?”
Strategy:
- Label IV / DV / design in the margin of your scratch paper.
- If no manipulation or random assignment → no firm causal claim.
- If results are association-only → answer with correlation language.
- Prefer operational clarity over jargon-heavy but empty options.
- Ethics wrong answers often hide coercion or missing debrief after deception.
Research methods is the “immune system” of Social Science: it keeps you from accepting every confident claim about mind, society, and health behavior. Master the table of designs, the reliability/validity contrast, and the ethics checklist, and many items become pattern recognition rather than guesswork.
A researcher randomly assigns students to two study methods and later compares exam scores. Exam score is best described as the:
Which conclusion is justified by a strong positive correlation between hours of sleep and quiz scores in an observational study?
Reliability of a psychological scale primarily refers to:
Informed consent in research ethics most directly requires that participants: