8.4 Biostatistics and Research Methods
Key Takeaways
- The hierarchy of evidence places systematic reviews and meta-analyses of RCTs at the top, followed by RCTs, cohort studies, case-control studies, case series, and expert opinion.
- A 95% confidence interval means that over repeated samples 95% of constructed intervals will contain the true parameter; if the interval crosses the null value, the result is not statistically significant at α = 0.05.
- Type I error (α) is a false positive and is typically set at 0.05; Type II error (β) is a false negative and power equals 1 − β.
- Intention-to-treat analysis preserves randomization benefits by analyzing patients in their assigned group regardless of adherence, while per-protocol analysis reflects efficacy under ideal conditions.
- The Belmont Report principles are respect for persons, beneficence, and justice; the Declaration of Helsinki provides international ethical guidance for medical research.
Biostatistics and Research Methods
Quick Answer: Biostatistics provides the tools to design studies, summarize data, and draw valid inferences. The FPGEE emphasizes study-design hierarchy, descriptive and inferential statistics, confidence intervals, statistical-test selection by data type, survival analysis, and research ethics grounded in the Belmont Report and Declaration of Helsinki.
Study Designs
Research designs divide into observational (no treatment assignment) and experimental (investigator assigns treatment).
- Cross-sectional — exposure and outcome measured at one time point; useful for prevalence.
- Cohort — follows exposed versus unexposed forward in time; supports causality and incidence.
- Case-control — compares exposure history in cases versus controls; ideal for rare diseases.
- Randomized controlled trial (RCT) — gold standard; randomization balances confounders and supports causal inference.
- Quasi-experimental — intervention without randomization (e.g., before/after policy changes); weaker than RCT.
Hierarchy of Evidence
Evidence pyramids rank designs by internal validity, from systematic reviews and meta-analyses of RCTs at the top, down through RCTs, cohort studies, case-control studies, case series and case reports, to expert opinion and animal studies at the base. Meta-analysis sits atop because it pools RCT results, increasing precision and detecting effects missed by single trials.
Descriptive Statistics
Descriptive statistics summarize a sample without inferring beyond it.
- Mean — arithmetic average; sensitive to outliers.
- Median — middle value; robust for skewed data.
- Mode — most frequent value.
- Standard deviation (SD) — average deviation from the mean; approximately 68% of values fall within ±1 SD in a normal distribution.
- Variance — SD squared.
- Interquartile range (IQR) — range between the 25th and 75th percentiles; preferred for skewed distributions.
Inferential Statistics
Inference generalizes from sample to population. The null hypothesis (H₀) asserts no effect; the alternative hypothesis (H₁) asserts an effect. The p-value is the probability of results at least as extreme as observed if H₀ is true. A p < 0.05 threshold rejects H₀ but does not measure effect size or clinical importance.
Type I and Type II Errors, Power
| Decision | H₀ True | H₀ False |
|---|---|---|
| Reject H₀ | Type I error (α) | Correct (power) |
| Fail to reject H₀ | Correct | Type II error (β) |
- Type I (α) — false positive; conventionally set at 0.05.
- Type II (β) — false negative; typically limited to 0.20.
- Power = 1 − β — probability of detecting a true effect; commonly targeted at 0.80. Power increases with sample size, effect size, and α.
Confidence Intervals
A 95% confidence interval means that, over repeated samples, 95% of intervals constructed will contain the true parameter. It is not the probability that the true value lies in this specific interval. Confidence intervals convey precision: a narrow CI suggests a precise estimate; a wide CI suggests more data are needed. For a relative risk, a 95% CI that crosses 1.0 indicates the result is not statistically significant at α = 0.05.
Statistical Test Selection
| Test | Data Type | Comparison |
|---|---|---|
| Paired t-test | Continuous, paired | Pre/post or matched pairs |
| Unpaired (two-sample) t-test | Continuous, independent | Two group means |
| One-way ANOVA | Continuous, independent | Three or more group means |
| Chi-square | Categorical, independent | Two or more proportions |
| Fisher's exact | Categorical, small counts | Rare events or cells under 5 |
| McNemar | Categorical, paired | Paired binary outcomes |
| Mann-Whitney U | Ordinal, independent | Two groups, non-normal |
| Kruskal-Wallis | Ordinal, independent | Three or more groups, non-normal |
For normally distributed continuous data, use parametric tests (t-test, ANOVA). For skewed or ordinal data, use non-parametric analogs (Mann-Whitney, Kruskal-Wallis).
Correlation and Regression
- Pearson correlation (r) — linear association between two continuous variables; range −1 to +1.
- Spearman correlation (ρ) — rank-based; preferred for ordinal or non-linear monotonic relationships.
- Linear regression — models a continuous outcome as a linear function of one (simple) or more (multiple) predictors.
- Logistic regression — models a binary outcome (e.g., response vs. no response); outputs odds ratios.
Survival Analysis
Kaplan-Meier estimation calculates survival probability over time, accounting for censored observations. The log-rank test compares survival curves between groups. Cox proportional hazards regression estimates a hazard ratio (HR) — the relative rate of events between groups — assuming hazards remain proportional over time.
Trial Analysis Approaches
- Intention-to-treat (ITT) — analyzes patients in the group to which they were randomized, regardless of adherence. Preserves randomization and reflects real-world effectiveness.
- Per-protocol — analyzes only patients who completed assigned treatment; more reflective of efficacy under ideal conditions but vulnerable to bias.
- Non-inferiority trials — seek to show a new treatment is not worse than the standard by more than a pre-specified margin (Δ). The CI upper bound must lie below Δ. They are typically one-sided and require different sample-size logic from superiority trials.
Bias Types
- Selection bias — systematic differences in how groups are chosen; mitigated by randomization.
- Information bias — measurement errors (recall, observer) that distort exposure-outcome relationships.
- Confounding — a third variable distorts the apparent association; addressed by randomization, stratification, or multivariable adjustment.
- Publication bias — studies with positive results are more likely to be published, distorting the literature; detected via funnel plots.
Data Collection and Research Ethics
Surveys and questionnaires must be validated (reliability — Cronbach's α ≥ 0.7; validity — content, construct, criterion). Research involving human subjects requires Institutional Review Board (IRB) approval and informed consent. The Belmont Report articulates three principles: respect for persons (autonomy, consent), beneficence (maximize benefit, minimize harm), and justice (fair subject selection). The Declaration of Helsinki (World Medical Association) provides international ethical guidelines emphasizing patient welfare over scientific interests. Pharmacists serving as investigators or co-investigators must complete Good Clinical Practice (GCP) training and follow 21 CFR Part 11 for electronic records.
Which statistical test is most appropriate for comparing three independent group means when the data are normally distributed?
A study reports a relative risk of 0.8 with a 95% confidence interval of 0.6 to 1.2. The most appropriate interpretation is:
Analyzing patients in the group to which they were randomized, regardless of adherence, is called:
The Belmont Report principle that requires fair selection of research subjects is: