15.3 Basic and Applied Research Design: Quantitative, Qualitative, Reliability, and Validity

Key Takeaways

  • Basic research generates generalizable knowledge; applied research answers practice-relevant questions; evidence-based practice integrates the best research evidence with clinical expertise and client values
  • Quantitative designs span experimental (RCT), quasi-experimental (nonequivalent control group, time series), descriptive/survey, and correlational; qualitative designs include phenomenology, grounded theory, ethnography, case study, narrative, and participatory/action research
  • Reliability is the consistency of a measure (test-retest, inter-rater, internal consistency); validity is whether it measures what it claims (content, criterion, construct, face) — a measure can be reliable without being valid but cannot be valid without being reliable
  • Threats to internal validity include history, maturation, testing, instrumentation, statistical regression to the mean, selection, and attrition; threats to external validity limit generalizability to other people, settings, and times
  • Statistical significance (p < .05) means a result is unlikely due to chance; clinical significance means the result is large enough to matter to a client — both are required for an evidence base, and exam vignettes distinguish them
Last updated: August 2026

Clinical social workers are both consumers and producers of research. The ASWB Clinical exam expects you to recognize research designs, evaluate whether a study supports its conclusions, and understand how evidence-based practice integrates research with client values and clinical expertise. This section is concept-heavy; the exam tests recognition and application, not computation.

Why Research Design Matters for the Exam

The exam uses research-design items in two ways. First, it gives you a study description and asks what type of design it is (experimental, quasi-experimental, qualitative, mixed). Second, it gives you a finding and asks whether the conclusion is warranted given the design's threats to validity. Knowing the canonical threats — history, maturation, testing, instrumentation, regression to the mean, selection, attrition — lets you eliminate wrong answers quickly.

Basic Versus Applied Research

Basic research aims to generate generalizable knowledge about human behavior, often without an immediate application. Applied research aims to answer questions that bear directly on practice, program, or policy. The two sit on a continuum; many studies have both basic and applied features. For a clinical social worker, the most relevant applied research is intervention research — studies that test whether a specific intervention, delivered in a specific way, produces change.

Quantitative Research Designs

Quantitative designs produce numerical data and emphasize control, measurement, and generalizability.

DesignDefining featureInternal validityTypical use
Experimental (RCT)Random assignment to treatment and control/comparison conditionsStrongestTesting causal effects of an intervention
Quasi-experimental (nonequivalent control group)Comparison group but no random assignmentModerate — selection bias is a key threatField settings where randomization is not feasible
Quasi-experimental (time series)Repeated measurement over time before and after an interventionModerate — history is a key threatEvaluating a policy change at the agency or community level
Descriptive / surveyCharacterizes a population or phenomenon at one point in timeWeak causal inference; strong descriptive valueNeeds assessment, prevalence
CorrelationalExamines association between variables without manipulationCannot support causationHypothesis generation

Random assignment is the feature that distinguishes a true experiment from a quasi-experiment. Without randomization, selection bias is always a threat: the groups may have differed before the intervention in ways that explain the post-intervention difference.

Qualitative Research Designs

Qualitative designs produce non-numerical data (words, observations, artifacts) and emphasize meaning, context, and depth.

DesignCentral questionProduct
PhenomenologyWhat is the lived experience of a phenomenon?Description of the essence of the experience
Grounded theoryWhat process or theory explains the phenomenon?A theory generated from the data
EthnographyWhat is the culture of this group or setting?A cultural description
Case studyWhat is happening in this single bounded case?In-depth case analysis
Narrative researchWhat does the participant's own story reveal?A coherent narrative account
Participatory / action researchHow can participants collaboratively study and change their own situation?Both knowledge and action

Qualitative sampling is typically purposive — participants are chosen because they can inform the question, not because they represent a population statistically. Sample size is determined by saturation — the point at which new interviews or observations stop producing new themes.

Mixed Methods

Mixed-methods research combines quantitative and qualitative data in one study. A common design is an explanatory sequential design: collect quantitative data first, then use qualitative interviews to explain and elaborate the quantitative findings. Exploratory sequential design reverses the order: qualitative first to develop a measure or theory, then quantitative to test it. Mixed methods are especially suited to program evaluation because they answer whether (quantitative) and why (qualitative) at once.

Sampling

For quantitative research, the goal is a representative sample drawn from a defined population. Key methods include:

  • Simple random sampling — every member of the population has an equal chance of selection
  • Stratified random sampling — population divided into strata (e.g., by race, gender), then random samples drawn from each stratum to ensure representation
  • Cluster sampling — randomly selected clusters (e.g., schools, agencies) rather than individuals
  • Convenience sampling — participants who are readily available; weak representativeness

Sample size for quantitative studies is driven by a power analysis that balances effect size, alpha (typically .05), and power (typically .80). Underpowered studies fail to detect real effects; overpowered studies detect trivially small effects as statistically significant.

Reliability and Validity of Measurement

Reliability is the consistency of a measure. Validity is whether it measures what it claims to measure. A measure can be reliable without being valid (it consistently measures the wrong thing), but it cannot be valid without being reliable.

Reliability typeWhat it measuresMethod
Test-retestStability over timeSame instrument to same people at two times; correlate
Inter-raterConsistency across ratersTwo or more raters score the same cases; compute agreement or kappa
Internal consistencyConsistency across items in a scaleCronbach's alpha (typically ≥ .70 acceptable)
Validity typeWhat it asksMethod
Face validityDoes it look like it measures the construct?Expert judgment, surface inspection
Content validityDoes it cover the full content domain?Expert review of item coverage
Criterion validityDoes it predict or agree with a gold standard?Concurrent or predictive correlation with a criterion
Construct validityDoes it measure the theoretical construct?Convergent and discriminant validity; factor analysis

Threats to Internal and External Validity

Internal validity is the degree to which a study supports a causal conclusion — whether the intervention, rather than a rival explanation, produced the change. External validity is the degree to which findings generalize beyond the study.

Threat (internal)Description
HistoryAn outside event during the study explains the change
MaturationParticipants naturally change over time (e.g., grow, recover)
TestingTaking a pretest changes performance on the posttest
InstrumentationThe measure itself changes (new rater, recalibrated tool)
Statistical regression to the meanExtreme scorers tend to score closer to the mean on retest
SelectionGroups differed before the intervention (no randomization)
AttritionDifferential dropout changes the composition of groups

Key threats to external validity include non-representative samples (do findings apply to other clients?), unique setting features (do they apply in other settings?), and reactivity (does being studied change behavior — the Hawthorne effect)?

Statistical Versus Clinical Significance

A result can be statistically significant (unlikely due to chance, p < .05) but clinically trivial (a 1-point drop on a 100-point depression scale). Conversely, a result can be clinically meaningful (clients return to work, resume parenting) but fail to reach statistical significance in an underpowered study. The exam rewards answers that distinguish the two and treat both as necessary for an evidence base. Effect-size measures (Cohen's d, odds ratios) bridge the gap by quantifying the magnitude of change.

Evidence-Based Practice as Integration

The Sackett definition of evidence-based practice, adapted for social work, integrates three streams:

  1. Best research evidence — including but not limited to RCTs; for many practice questions, single-case designs, qualitative studies, and systematic reviews all contribute
  2. Clinical expertise — the social worker's judgment, formed by training and experience
  3. Client values, preferences, and circumstances — including culture, context, and expressed goals

No single stream is sufficient. The exam often presents distractors that privilege research evidence at the expense of client values ("apply the protocol regardless of the client's preference") or privilege client values at the expense of evidence ("do whatever the client asks"). The correct answer integrates all three.

Reading and Evaluating Research Critically

When evaluating a study, ask:

  • Was the sample adequate and representative? Underpowered or convenience samples limit conclusions.
  • Was the design appropriate to the question? A correlational design cannot answer a causal question.
  • Were the measures reliable and valid? A study using an unvalidated scale cannot support strong conclusions.
  • Were threats to internal validity addressed? Randomization, a comparison group, and pre/post measurement all strengthen causal inference.
  • Is the finding clinically meaningful, not just statistically significant? Report effect sizes, not just p-values.
  • Were ethical safeguards in place? Informed consent, IRB approval, confidentiality protections (the bridge to research ethics, covered in Chapter 5).
Test Your Knowledge

A researcher assigns half of a community mental health agency's clinicians to training in a new evidence-based treatment and compares client outcomes between trained and untrained clinicians. Clinicians volunteer for the training rather than being randomly assigned. Which threat to internal validity is most salient?

A
B
C
D
Relative exam emphasis among threats to internal validity (illustrative)
Test Your Knowledge

A researcher develops a new scale to measure therapeutic alliance. Cronbach's alpha for the scale is .88 in the pilot sample, and the scale correlates strongly (r = .62) with an established alliance measure. Which statement is MOST accurate?

A
B
C
D
Test Your Knowledge

A randomized controlled trial of a brief intervention for depression in older adults reports a statistically significant reduction in PHQ-9 scores (p = .04) but a mean difference of only 1.2 points between groups. Which interpretation is MOST consistent with the distinction between statistical and clinical significance?

A
B
C
D