14.4 Research Design, Statistics, and Program Evaluation
Key Takeaways
- True experimental designs establish causality through random assignment (R), independent variable (IV) manipulation, and control groups, whereas quasi-experimental designs lack randomization and pre-experimental designs lack comparison groups.
- Internal validity assesses whether the experimental manipulation (IV) solely caused the observed change in the dependent variable (DV), threatened by history, maturation, testing, instrumentation, regression, selection bias, and attrition.
- External validity assesses generalizability, threatened by participant reactivity (Hawthorne effect) and experimenter expectancy (Rosenthal/Pygmalion effect), mitigated through double-blind research protocols.
- In inferential hypothesis testing, Type I error (alpha) represents a false positive rejection of a true null hypothesis, while Type II error (beta) represents a false negative retention of a false null hypothesis; statistical power is defined as 1 - beta.
- Program evaluation distinguishes formative evaluation (ongoing process monitoring and implementation fidelity) from summative evaluation (terminal outcome and impact assessment), operationalized through Daniel Stufflebeam's CIPP model.
14.4 Research Design, Statistics, and Program Evaluation
Quick Answer: Research competence is an ethical mandate (CACREP Domain 8) ensuring counselors consume, evaluate, and produce evidence-based clinical interventions. Research designs exist on a continuum of internal validity from true experimental (random assignment, manipulated IV, control group) through quasi-experimental (intact groups, no randomization) to pre-experimental and qualitative methodologies. Threats to internal validity (history, maturation, regression to the mean, attrition) and external validity (Hawthorne effect, Rosenthal effect) must be controlled. In statistics, clinicians must master the normal distribution (68–95–99.7 rule), measures of central tendency (mean, median, mode), and inferential decision-making (Type I error [$\alpha$] vs. Type II error [$\beta$], power [$1 - \beta$], $t$-tests, ANOVA, Chi-Square, Pearson $r$). Program evaluation bridges research and clinical practice via formative vs. summative designs and Stufflebeam's CIPP model.
Quantitative vs. Qualitative Research Paradigms
Clinical research is organized around two overarching philosophical paradigms:
- Quantitative Research Paradigm: Grounded in post-positivist epistemology. Assumes an objective, measurable reality that can be deconstructed into discrete operational variables. Employs deductive logic ($Theory \to Hypothesis \to Observation \to Confirmation$), standardized psychometric instruments, numerical data, statistical hypothesis testing, and tight experimental controls to establish generalizable causal relationships.
- Qualitative Research Paradigm: Grounded in constructivist, interpretivist, and phenomenological epistemologies. Assumes multiple, socially constructed realities. Employs inductive logic ($Observation \to Pattern \to Tentative \ Hypothesis \to Theory$), naturalistic field observation, semi-structured interviews, thematic analysis, and "thick description." The researcher acts as the primary data-collection instrument (e.g., Grounded Theory, Phenomenology, Ethnography, Case Studies).
Experimental Research Design Continuum
Donald Campbell and Julian Stanley established the definitive taxonomy structuring quantitative research into three distinct levels of experimental rigor, defined by their capacity to isolate causal relationships:
[ Pre-Experimental ] ──────> [ Quasi-Experimental ] ──────> [ True Experimental ]
(No Randomization, (No Randomization, (Random Assignment [R],
No Control Group) Intact Comparison Groups) Active IV, Control Group)
*Lowest Internal Validity* *Highest Internal Validity*
1. True Experimental Designs (Gold Standard for Causality)
True experiments possess three non-negotiable criteria:
- Random Assignment ($R$): Every participant has an equal, independent probability of being assigned to either the experimental group or the control group. Random assignment neutralizes systematic participant differences prior to intervention.
- Active Manipulation of the Independent Variable ($IV$): The researcher actively introduces, controls, and manipulates the experimental treatment ($IV$) while measuring its effect on the outcome variable ($DV$).
- Presence of a Control or Comparison Group: An untreated control, placebo, or treatment-as-usual comparison group to benchmark change.
- Classic Designs: Pretest-Posttest Control Group Design, Posttest-Only Control Group Design, Solomon Four-Group Design (the premier design that tests and controls for pretest sensitization).
2. Quasi-Experimental Designs
Quasi-experiments actively manipulate the Independent Variable and include comparison groups, but lack random assignment. Researchers utilize intact, pre-existing cohorts (e.g., comparing two existing classrooms, two hospital wards, or two community clinics).
- Classic Designs: Non-equivalent Comparison Group Design, Interrupted Time-Series Design.
- Clinical Reality: Common in counseling field research where withholding treatment or randomly reassigning clients across clinics is ethically or logistically impossible. However, selection bias remains an inherent threat.
3. Pre-Experimental Designs
Pre-experimental designs lack both random assignment and a proper control group. They exhibit extremely weak internal validity and cannot establish causality.
- Classic Designs: One-Shot Case Study ($X \ O$), One-Group Pretest-Posttest Design ($O_1 \ X \ O_2$). Changes between pretest and posttest may easily stem from historical events or natural maturation rather than the intervention.
Single-Case Experimental Designs (SCED)
Single-subject research evaluates treatment efficacy on an individual client or small group across time by using the client as their own control:
- $AB$ Design: Baseline phase ($A$) followed by treatment phase ($B$). Weakest design; cannot rule out history.
- $ABAB$ Reversal Design: Baseline ($A$), Treatment ($B$), Withdrawal of treatment ($A$), Reintroduction of treatment ($B$). If target behavior fluctuates exclusively when the treatment is present, experimental causality is demonstrated.
- Ethical Dilemma: Withdrawing an effective intervention ($A$) is ethically contraindicated if the client presents with acute self-harm, suicidal ideation, or severe aggressive behaviors.
Threats to Internal and External Validity
Threats to Internal Validity (Campbell & Stanley)
Internal validity represents the degree of confidence that the manipulation of the Independent Variable ($IV$) solely produced the observed change in the Dependent Variable ($DV$), free from confounding variables:
| Threat | Clinical Mechanism & Operational Definition |
|---|---|
| 1. History | Specific, unplanned external events occurring between the pretest and posttest that influence the DV (e.g., an economic recession, a campus tragedy, or new federal policy during a study on student anxiety). |
| 2. Maturation | Biological, physiological, or psychological growth occurring within participants purely as a function of the passage of time (e.g., clients naturally healing from acute grief, children developing cognitive capacity). |
| 3. Testing | The practice effect or sensitization resulting from completing a pretest, which alters performance on subsequent posttests independent of any treatment effect. |
| 4. Instrumentation | Changes in the calibration of measuring instruments, observer fatigue, diagnostic drift, or altering scoring rubrics between pretest and posttest administrations. |
| 5. Statistical Regression | The mathematical tendency for extremely high or extremely low scores to naturally drift (regress) toward the group mean upon retesting due to measurement error. |
| 6. Selection Bias | Systematic differences in participant characteristics between groups prior to intervention, resulting from a lack of true random assignment. |
| 7. Attrition / Mortality | The differential, non-random loss of participants from experimental versus control groups over the course of an investigation, skewing final group comparability. |
Threats to External Validity (Generalizability)
External validity represents the degree to which experimental findings can be generalized across diverse populations, settings, and times:
- Hawthorne Effect (Reactivity): Participants alter their natural behavior simply because they are aware they are being observed and evaluated in an experiment.
- Rosenthal Effect (Pygmalion / Experimenter Expectancy): The researcher's preconceived biases, hopes, or expectations unconsciously influence participant performance or observer scoring. Mitigated through double-blind protocols where neither the participant nor the administering clinician knows who receives active treatment.
- Reactive Effects of Testing (Pretest Sensitization): Completing a pretest increases participant sensitivity or responsiveness to the experimental intervention, meaning findings cannot be generalized to an un-pretested population.
Descriptive Statistics and the Normal Distribution
Scales of Measurement: The NOIR Framework
- Nominal: Categorical, qualitative grouping with no mathematical order or magnitude (e.g., gender, race, diagnostic categories, yes/no). Mode is the only legitimate measure of central tendency.
- Ordinal: Ranked order without equal intervals between ranks (e.g., class rank, Likert scales, socioeconomic strata). Median is the preferred measure.
- Interval: Ordered scale with equal, standardized intervals between units, but lacks an absolute, true zero point (e.g., Fahrenheit/Celsius temperature, standard IQ test scores). Mean and Standard Deviation are legitimate.
- Ratio: Equal intervals with a true, absolute zero point representing the complete absence of the property (e.g., weight, height, annual income, response latency in seconds). All mathematical operations permitted.
Measures of Central Tendency and Dispersion
- Central Tendency:
- Mean: Arithmetic average. Highly sensitive to extreme outliers; distorted in skewed distributions.
- Median: The exact 50th percentile dividing a distribution in half. The preferred measure for skewed data (e.g., income, real estate prices).
- Mode: The most frequently occurring score in a distribution. Unimodal, bimodal, or multimodal.
- Measures of Dispersion (Variability):
- Range: Highest score minus lowest score ($+ 1$). Crude measure of spread.
- Variance ($s^2$ or $\sigma^2$): The average of squared deviations from the mean.
- Standard Deviation ($s$ or $\sigma$): The square root of variance. Expresses dispersion in the original units of measurement.
Normal Bell Curve
│
Mean
Median
Mode
┌────────────┴────────────┐
-1 SD +1 SD
[ <────────── 68.26% ──────────> ]
-2 SD +2 SD
[ <────────────────── 95.44% ──────────────────> ]
-3 SD +3 SD
[ <──────────────────────── 99.74% ────────────────────────> ]
The Normal Distribution (Gaussian Curve) and Empirical Rule
Under a theoretical standard normal distribution:
- Symmetrical, bell-shaped, unimodal distribution where Mean = Median = Mode at the exact center.
- The Empirical Rule (68–95–99.7 Rule):
- Exactly 68.26% of all scores fall within $\pm 1$ Standard Deviation of the mean ($34.13%$ on each side).
- Exactly 95.44% of all scores fall within $\pm 2$ Standard Deviations of the mean ($47.72%$ on each side).
- Exactly 99.74% of all scores fall within $\pm 3$ Standard Deviations of the mean ($49.87%$ on each side).
Distribution Skewness
- Positively Skewed (Skewed to the Right): Tail points toward the positive (right) end. A cluster of low scores with a few extreme high outliers pulls the mean upward. Mode < Median < Mean (e.g., national income distribution).
- Negatively Skewed (Skewed to the Left): Tail points toward the negative (left) end. A cluster of high scores with a few extreme low outliers pulls the mean downward. Mean < Median < Mode (e.g., a very easy examination where most students score 95%).
Inferential Statistics, Hypothesis Testing, and Decision Errors
Inferential statistics allow researchers to draw mathematical inferences about a broad population based on probability samples.
Hypothesis Testing and the Null Hypothesis
- Null Hypothesis ($H_0$): Posits that there is no true difference, no treatment effect, or no relationship between variables in the population; any observed difference is purely due to chance or sampling error.
- Alternative / Directional Hypothesis ($H_1$): Posits that a real difference, treatment effect, or relationship exists between variables.
- Alpha Level ($\alpha$): The critical significance threshold established a priori by the researcher (conventionally set at $\alpha = .05$ or $\alpha = .01$). If the obtained probability ($p$-value) is less than alpha ($p < .05$), the researcher rejects the null hypothesis, concluding that the finding is statistically significant.
The Type I vs. Type II Error Matrix
| Statistical Decision | True State of Reality: $H_0$ is True (No Effect) | True State of Reality: $H_0$ is False (Real Effect Exists) |
|---|---|---|
| Reject Null Hypothesis ($H_0$) | Type I Error ($\alpha$)<br>(False Positive: Claiming treatment works when it does not) | Correct Decision ($1 - \beta$)<br>(Statistical Power: Successfully detecting real effect) |
| Fail to Reject (Retain) $H_0$ | Correct Decision ($1 - \alpha$)<br>(Confidence Level: Accurately retaining true null) | Type II Error ($\beta$)<br>(False Negative: Missing a real treatment effect) |
- Type I Error (Alpha, $\alpha$): Rejecting a null hypothesis that is actually true. The researcher claims a treatment effect exists when it was purely random chance (a false positive). Set directly by the alpha level (at $\alpha = .05$, there is a 5% chance of Type I error).
- Type II Error (Beta, $\beta$): Failing to reject a null hypothesis that is actually false. The researcher concludes there is no treatment effect when a real clinical difference actually existed (a false negative).
- Statistical Power ($1 - \beta$): The probability of correctly rejecting a false null hypothesis (detecting a real effect when one exists). Power is increased by: (1) increasing sample size ($N$), (2) increasing the alpha level (e.g., from .01 to .05), (3) increasing effect size, and (4) utilizing directional (one-tailed) tests.
Major Statistical Significance Tests
| Statistical Test | Design Requirements | Variables & Assumptions |
|---|---|---|
| Independent Samples $t$-test | Compares means of two independent groups | 1 categorical IV with 2 levels (e.g., CBT vs. Control); 1 continuous DV (interval/ratio). Assumes normality and homogeneity of variance. |
| Paired (Dependent) $t$-test | Compares means of one group across two time points or matched pairs | Within-subjects design (e.g., Pretest mean vs. Posttest mean for the same participants). |
| One-Way ANOVA ($F$-ratio) | Compares means of three or more independent groups on one DV | 1 categorical IV with 3+ levels (e.g., CBT vs. Psychodynamic vs. Control); 1 continuous DV. Avoids inflating familywise Type I error. If $F$ is significant, requires post-hoc tests (Tukey's HSD, Scheffé) to pinpoint pairwise differences. |
| Two-Way ANOVA | Evaluates effects of two independent variables simultaneously | 2 categorical IVs (e.g., Therapy Type [CBT/Waitlist] $\times$ Gender [Male/Female]) on 1 continuous DV. Yields three distinct $F$-tests: Main effect of IV1, Main effect of IV2, and Interaction effect ($IV1 \times IV2$). |
| Chi-Square ($\chi^2$) | Nonparametric test for categorical / nominal data | • Goodness-of-Fit: Tests whether observed frequencies fit expected theoretical distribution across 1 variable.<br>• Test of Independence: Evaluates association between 2 categorical variables (e.g., Gender $\times$ Diagnostic Category). |
| Pearson Correlation ($r$) | Measures linear relationship between two continuous variables | Parametric correlation for interval/ratio data. Ranges from $-1.00$ to $+1.00$. Coefficient of Determination ($r^2$) measures shared variance (e.g., if $r = .60$, then $r^2 = .36$, meaning $36%$ of variance is shared). |
| Spearman Rho ($\rho$) | Nonparametric correlation for ranked data | Evaluates monotonic relationship between two ordinal variables. |
Statistical vs. Practical Significance (Effect Size)
Statistical significance ($p < .05$) merely indicates that an outcome is unlikely due to chance; with a massive sample size ($N = 10,000$), trivial differences become statistically significant. In contrast, practical significance evaluates the clinical magnitude of an intervention via effect size, independent of sample size:
- Cohen's $d$: Standardized mean difference ($d = 0.2$ is small; $d = 0.5$ is medium; $d = 0.8$ is large).
Program Evaluation Models and Needs Assessment
Counselors conduct program evaluation to determine the utility, fidelity, and impact of mental health programs within clinical agencies and schools.
Formative vs. Summative Evaluation
- Formative Evaluation (Process Evaluation): Ongoing evaluation conducted during program planning and implementation. Assesses whether activities are occurring as planned, tracks service delivery fidelity, and identifies operational bottlenecks to allow real-time adjustments ("When the chef tastes the soup, it is formative").
- Summative Evaluation (Outcome Evaluation): Terminal evaluation conducted after program completion. Evaluates whether overall program goals and client benchmarks were achieved, assessing efficacy, impact, and cost-effectiveness ("When the guests taste the soup, it is summative").
Daniel Stufflebeam's CIPP Model
A comprehensive, widely tested decision-oriented evaluation model structuring evaluation across four phases:
- Context Evaluation: Evaluates the overarching institutional environment, diagnoses unmet needs, and identifies systemic barriers to define target goals (needs assessment).
- Input Evaluation: Assesses available resources, staffing capabilities, agency budget, and competing strategic plans to select the most feasible intervention strategy.
- Process Evaluation: Monitors ongoing day-to-day program implementation, assessing procedural fidelity and operational execution.
- Product Evaluation: Measures, interprets, and judges final program outcomes, determining whether the program should be continued, expanded, modified, or terminated.
Needs Assessment Protocols
Needs assessments determine the gap between "what is" (current status) and "what should be" (desired future competency) through systematic data-gathering:
- Methodologies: Surveys, focus groups, community forums, key informant interviews, and analyzing archival epidemiological data.
On the NCE Exam: Common Pitfalls and Key Distinctions
- Type I vs. Type II Error: Type I is a False Positive (rejecting a true null; claiming an effect exists when it does not). Type II is a False Negative (retaining a false null; failing to detect a real effect).
- ANOVA vs. Multiple $t$-tests: Never run multiple independent $t$-tests to compare three or more groups! Running multiple $t$-tests exponentially inflates the experimentwise / familywise Type I error rate. Always run a One-Way ANOVA, followed by post-hoc tests (e.g., Tukey HSD) only if the omnibus $F$-test is significant.
- Coefficient of Determination ($r^2$): Always square the Pearson $r$ correlation coefficient to calculate the proportion of shared variance between two variables. If $r = .70$, the shared variance is $.49$ (or $49%$), not $70%$.
A counseling researcher is evaluating the efficacy of a novel virtual reality mindfulness intervention compared to traditional CBT and a waitlist control group in reducing generalized anxiety across a sample of 90 randomized participants. The researcher measures anxiety on an interval scale after eight weeks of treatment. To determine whether statistically significant differences exist among the three group means without inflating the experimentwise Type I error rate, which inferential statistical test should the researcher conduct?
A clinic director conducts an empirical outcome study testing whether an intensive trauma protocol decreases PTSD symptom severity. After statistical analysis, the director finds a p-value of .18 and fails to reject the null hypothesis, concluding that the trauma protocol produces no therapeutic benefit. However, in reality, the treatment is extraordinarily effective, but the study lacked an adequate sample size (N = 12) to achieve statistical significance. What methodological phenomenon has occurred?
A community mental health agency is designing an adolescent substance abuse prevention initiative. The program evaluator utilizes Daniel Stufflebeam's CIPP model. During the initial planning phase, the evaluator reviews regional demographic trends, surveys local school counselors, assesses community risk factors, and diagnoses the specific unmet prevention needs of the target population. Which component of the CIPP model is being executed?