15.3 Basic and Applied Research Design: Quantitative, Qualitative, Reliability, and Validity
Key Takeaways
- Basic research generates generalizable knowledge; applied research answers practice-relevant questions; evidence-based practice integrates the best research evidence with clinical expertise and client values
- Quantitative designs span experimental (RCT), quasi-experimental (nonequivalent control group, time series), descriptive/survey, and correlational; qualitative designs include phenomenology, grounded theory, ethnography, case study, narrative, and participatory/action research
- Reliability is the consistency of a measure (test-retest, inter-rater, internal consistency); validity is whether it measures what it claims (content, criterion, construct, face) — a measure can be reliable without being valid but cannot be valid without being reliable
- Threats to internal validity include history, maturation, testing, instrumentation, statistical regression to the mean, selection, and attrition; threats to external validity limit generalizability to other people, settings, and times
- Statistical significance (p < .05) means a result is unlikely due to chance; clinical significance means the result is large enough to matter to a client — both are required for an evidence base, and exam vignettes distinguish them
Clinical social workers are both consumers and producers of research. The ASWB Clinical exam expects you to recognize research designs, evaluate whether a study supports its conclusions, and understand how evidence-based practice integrates research with client values and clinical expertise. This section is concept-heavy; the exam tests recognition and application, not computation.
Why Research Design Matters for the Exam
The exam uses research-design items in two ways. First, it gives you a study description and asks what type of design it is (experimental, quasi-experimental, qualitative, mixed). Second, it gives you a finding and asks whether the conclusion is warranted given the design's threats to validity. Knowing the canonical threats — history, maturation, testing, instrumentation, regression to the mean, selection, attrition — lets you eliminate wrong answers quickly.
Basic Versus Applied Research
Basic research aims to generate generalizable knowledge about human behavior, often without an immediate application. Applied research aims to answer questions that bear directly on practice, program, or policy. The two sit on a continuum; many studies have both basic and applied features. For a clinical social worker, the most relevant applied research is intervention research — studies that test whether a specific intervention, delivered in a specific way, produces change.
Quantitative Research Designs
Quantitative designs produce numerical data and emphasize control, measurement, and generalizability.
| Design | Defining feature | Internal validity | Typical use |
|---|---|---|---|
| Experimental (RCT) | Random assignment to treatment and control/comparison conditions | Strongest | Testing causal effects of an intervention |
| Quasi-experimental (nonequivalent control group) | Comparison group but no random assignment | Moderate — selection bias is a key threat | Field settings where randomization is not feasible |
| Quasi-experimental (time series) | Repeated measurement over time before and after an intervention | Moderate — history is a key threat | Evaluating a policy change at the agency or community level |
| Descriptive / survey | Characterizes a population or phenomenon at one point in time | Weak causal inference; strong descriptive value | Needs assessment, prevalence |
| Correlational | Examines association between variables without manipulation | Cannot support causation | Hypothesis generation |
Random assignment is the feature that distinguishes a true experiment from a quasi-experiment. Without randomization, selection bias is always a threat: the groups may have differed before the intervention in ways that explain the post-intervention difference.
Qualitative Research Designs
Qualitative designs produce non-numerical data (words, observations, artifacts) and emphasize meaning, context, and depth.
| Design | Central question | Product |
|---|---|---|
| Phenomenology | What is the lived experience of a phenomenon? | Description of the essence of the experience |
| Grounded theory | What process or theory explains the phenomenon? | A theory generated from the data |
| Ethnography | What is the culture of this group or setting? | A cultural description |
| Case study | What is happening in this single bounded case? | In-depth case analysis |
| Narrative research | What does the participant's own story reveal? | A coherent narrative account |
| Participatory / action research | How can participants collaboratively study and change their own situation? | Both knowledge and action |
Qualitative sampling is typically purposive — participants are chosen because they can inform the question, not because they represent a population statistically. Sample size is determined by saturation — the point at which new interviews or observations stop producing new themes.
Mixed Methods
Mixed-methods research combines quantitative and qualitative data in one study. A common design is an explanatory sequential design: collect quantitative data first, then use qualitative interviews to explain and elaborate the quantitative findings. Exploratory sequential design reverses the order: qualitative first to develop a measure or theory, then quantitative to test it. Mixed methods are especially suited to program evaluation because they answer whether (quantitative) and why (qualitative) at once.
Sampling
For quantitative research, the goal is a representative sample drawn from a defined population. Key methods include:
- Simple random sampling — every member of the population has an equal chance of selection
- Stratified random sampling — population divided into strata (e.g., by race, gender), then random samples drawn from each stratum to ensure representation
- Cluster sampling — randomly selected clusters (e.g., schools, agencies) rather than individuals
- Convenience sampling — participants who are readily available; weak representativeness
Sample size for quantitative studies is driven by a power analysis that balances effect size, alpha (typically .05), and power (typically .80). Underpowered studies fail to detect real effects; overpowered studies detect trivially small effects as statistically significant.
Reliability and Validity of Measurement
Reliability is the consistency of a measure. Validity is whether it measures what it claims to measure. A measure can be reliable without being valid (it consistently measures the wrong thing), but it cannot be valid without being reliable.
| Reliability type | What it measures | Method |
|---|---|---|
| Test-retest | Stability over time | Same instrument to same people at two times; correlate |
| Inter-rater | Consistency across raters | Two or more raters score the same cases; compute agreement or kappa |
| Internal consistency | Consistency across items in a scale | Cronbach's alpha (typically ≥ .70 acceptable) |
| Validity type | What it asks | Method |
|---|---|---|
| Face validity | Does it look like it measures the construct? | Expert judgment, surface inspection |
| Content validity | Does it cover the full content domain? | Expert review of item coverage |
| Criterion validity | Does it predict or agree with a gold standard? | Concurrent or predictive correlation with a criterion |
| Construct validity | Does it measure the theoretical construct? | Convergent and discriminant validity; factor analysis |
Threats to Internal and External Validity
Internal validity is the degree to which a study supports a causal conclusion — whether the intervention, rather than a rival explanation, produced the change. External validity is the degree to which findings generalize beyond the study.
| Threat (internal) | Description |
|---|---|
| History | An outside event during the study explains the change |
| Maturation | Participants naturally change over time (e.g., grow, recover) |
| Testing | Taking a pretest changes performance on the posttest |
| Instrumentation | The measure itself changes (new rater, recalibrated tool) |
| Statistical regression to the mean | Extreme scorers tend to score closer to the mean on retest |
| Selection | Groups differed before the intervention (no randomization) |
| Attrition | Differential dropout changes the composition of groups |
Key threats to external validity include non-representative samples (do findings apply to other clients?), unique setting features (do they apply in other settings?), and reactivity (does being studied change behavior — the Hawthorne effect)?
Statistical Versus Clinical Significance
A result can be statistically significant (unlikely due to chance, p < .05) but clinically trivial (a 1-point drop on a 100-point depression scale). Conversely, a result can be clinically meaningful (clients return to work, resume parenting) but fail to reach statistical significance in an underpowered study. The exam rewards answers that distinguish the two and treat both as necessary for an evidence base. Effect-size measures (Cohen's d, odds ratios) bridge the gap by quantifying the magnitude of change.
Evidence-Based Practice as Integration
The Sackett definition of evidence-based practice, adapted for social work, integrates three streams:
- Best research evidence — including but not limited to RCTs; for many practice questions, single-case designs, qualitative studies, and systematic reviews all contribute
- Clinical expertise — the social worker's judgment, formed by training and experience
- Client values, preferences, and circumstances — including culture, context, and expressed goals
No single stream is sufficient. The exam often presents distractors that privilege research evidence at the expense of client values ("apply the protocol regardless of the client's preference") or privilege client values at the expense of evidence ("do whatever the client asks"). The correct answer integrates all three.
Reading and Evaluating Research Critically
When evaluating a study, ask:
- Was the sample adequate and representative? Underpowered or convenience samples limit conclusions.
- Was the design appropriate to the question? A correlational design cannot answer a causal question.
- Were the measures reliable and valid? A study using an unvalidated scale cannot support strong conclusions.
- Were threats to internal validity addressed? Randomization, a comparison group, and pre/post measurement all strengthen causal inference.
- Is the finding clinically meaningful, not just statistically significant? Report effect sizes, not just p-values.
- Were ethical safeguards in place? Informed consent, IRB approval, confidentiality protections (the bridge to research ethics, covered in Chapter 5).
A researcher assigns half of a community mental health agency's clinicians to training in a new evidence-based treatment and compares client outcomes between trained and untrained clinicians. Clinicians volunteer for the training rather than being randomly assigned. Which threat to internal validity is most salient?
A researcher develops a new scale to measure therapeutic alliance. Cronbach's alpha for the scale is .88 in the pilot sample, and the scale correlates strongly (r = .62) with an established alliance measure. Which statement is MOST accurate?
A randomized controlled trial of a brief intervention for depression in older adults reports a statistically significant reduction in PHQ-9 scores (p = .04) but a mean difference of only 1.2 points between groups. Which interpretation is MOST consistent with the distinction between statistical and clinical significance?