12.3 Research Methods in Social Science

Key Takeaways

  • Independent variables are manipulated (or designated predictors); dependent variables are measured outcomes—confounds threaten causal claims.
  • Experiments support cause–effect when random assignment and control are strong; correlational designs show association only (correlation ≠ causation).
  • Reliability is consistency of measurement; validity is whether the measure captures the intended construct and supports intended inferences.
  • Ethics require informed consent, risk minimization, and careful limits on deception; NMAT tests research literacy, not advanced statistics software skills.
Last updated: August 2026

12.3 Research Methods in Social Science

Quick Answer: Know what was manipulated (IV), what was measured (DV), whether the design was experimental or correlational, whether the sample supports generalization, whether measures are reliable and valid, and whether ethics were respected. Correlation never proves causation.

CEM’s Social Science block rewards scientific thinking about people—the same literacy that protects you from weak claims in media, advertising, and even early research readings in medical school. You do not need advanced regression modeling; you need clean concepts.

Variables and Hypotheses

A hypothesis is a testable prediction. Variables are measurable attributes:

TermDefinitionExample
Independent variable (IV)Factor manipulated or treated as predictorStudy method (flashcards vs. rereading)
Dependent variable (DV)Outcome measuredQuiz score
Confounding variableUncontrolled factor that covaries with IV and could explain DVPrior GPA differing by condition
Operational definitionConcrete procedure for measuring a construct“Anxiety” = score on a named scale

Control variables are held constant or statistically adjusted so they do not obscure the IV–DV link. Random assignment (experiments) balances unknown confounds across groups on average; random sampling (from a population) supports external validity (generalization). Do not confuse the two.

Experimental vs. Correlational Designs

Experiments

In a true experiment, the researcher manipulates at least one IV, uses control conditions, and ideally randomly assigns participants. When these conditions hold, causal inference is strongest: changes in the IV cause changes in the DV (within design limits).

  • Between-subjects: different people in each condition.
  • Within-subjects: same people experience multiple conditions (order effects need counterbalancing).
  • Placebo / single- and double-blind procedures reduce expectancy and observer bias when applicable.

Correlational and descriptive studies

Correlational research measures naturally occurring variables without manipulation. A correlation coefficient (r) indexes direction and strength of linear association (−1 to +1). Positive correlation: both rise together; negative: one rises as the other falls; near zero: little linear association.

Correlation ≠ causation. Three classic threats:

  1. Directionality: Does A cause B or B cause A?
  2. Third variable: C causes both A and B.
  3. Spurious / chance associations in noisy data.

Descriptive methods (case study, naturalistic observation, surveys) map phenomena and generate hypotheses but rarely alone establish cause.

DesignManipulation?Best claim strength
True experimentYes + random assignmentCausal (with caveats)
Quasi-experimentIV-like groups without full random assignmentWeaker causal
CorrelationalNoAssociation only
DescriptiveNoDescription / exploration

Sampling

A population is the full set of interest; a sample is the studied subset.

  • Probability sampling (simple random, stratified, systematic) gives each unit a known chance of selection—supports statistical generalization when executed well.
  • Nonprobability sampling (convenience, snowball, voluntary response) is common and useful for early work but limits external validity.
  • Sample size affects precision and statistical power; larger is not automatically “more valid” if the sample is systematically biased.
  • Volunteer bias and undercoverage of hard-to-reach groups matter in health surveys.

Reliability and Validity

ConceptQuestion answeredExamples / subtypes
ReliabilityIs the measure consistent?Test–retest, inter-rater, internal consistency
ValidityDoes it measure what it claims and support intended uses?Content, criterion, construct; internal vs. external validity of studies

A scale can be reliable but not valid (consistently wrong construct). A study can have strong internal validity (tight causal control) but weak external validity (artificial lab). NMAT stems often ask which threat applies (confound, nonrandom sample, demand characteristics).

Demand characteristics are cues that lead participants to guess the hypothesis and alter behavior. Social desirability bias skews self-reports toward looking good. Observer bias distorts coding when raters know expected outcomes—hence blind coding and structured instruments.

Ethics in Human Research

Core principles (aligned with modern institutional review norms):

  1. Respect for persons: voluntary participation; informed consent with understandable information about procedures, risks, benefits, and withdrawal rights.
  2. Beneficence: minimize harm; risk–benefit balance; protect confidentiality/privacy of data.
  3. Justice: fair participant selection; avoid exploiting vulnerable groups without justification and safeguards.

Deception is sometimes used when full foreknowledge would destroy the study, but it is limited: scientific necessity, no unreasonable harm, and typically debriefing afterward. Historical abuses (and even well-known classic social psychology studies) drove today’s stricter oversight. NMAT may ask what ethical step was missing (no consent, no debrief, coercion).

Descriptive Statistics (Conceptual)

MeasureMeaningWhen it shines / fails
MeanArithmetic averageSensitive to extreme outliers
MedianMiddle valueRobust to outliers; skewed income-like data
ModeMost frequent valueCategories; multimodal distributions
Range / variance / SDSpreadHigh SD → more variability around mean

Know that a skewed distribution pulls the mean toward the tail more than the median. Bar graphs vs. histograms, and reading axes, are quantitative-transfer skills; Social Science items stay conceptual.

Statistical significance (e.g., p < .05 culture) means results are unlikely under a null model—not that the effect is large or important. Effect size and practical significance are separate ideas; if an option confuses “significant” with “huge real-world impact,” reject it.

How NMAT-Style Items Test Research Literacy

Expect stems such as:

  • “Which variable is independent?”
  • “Why can the researcher not conclude causation?”
  • “Which sampling problem is illustrated?”
  • “Which ethical principle was violated?”
  • “Which measure of central tendency is least affected by an extreme score?”

Strategy:

  1. Label IV / DV / design in the margin of your scratch paper.
  2. If no manipulation or random assignment → no firm causal claim.
  3. If results are association-only → answer with correlation language.
  4. Prefer operational clarity over jargon-heavy but empty options.
  5. Ethics wrong answers often hide coercion or missing debrief after deception.

Research methods is the “immune system” of Social Science: it keeps you from accepting every confident claim about mind, society, and health behavior. Master the table of designs, the reliability/validity contrast, and the ethics checklist, and many items become pattern recognition rather than guesswork.

Test Your Knowledge

A researcher randomly assigns students to two study methods and later compares exam scores. Exam score is best described as the:

A
B
C
D
Test Your Knowledge

Which conclusion is justified by a strong positive correlation between hours of sleep and quiz scores in an observational study?

A
B
C
D
Test Your Knowledge

Reliability of a psychological scale primarily refers to:

A
B
C
D
Test Your Knowledge

Informed consent in research ethics most directly requires that participants:

A
B
C
D