15.4 Biostatistics: Sampling, Central Tendency & Significance Testing
Key Takeaways
- Quantitative continuous data are summarized by Mean ± SD when normally distributed, and Median (IQR) when skewed.
- In a Normal (Gaussian) distribution: Mean ± 1 SD covers 68.3%, Mean ± 2 SD covers 95.4% (approximated as Mean ± 1.96 SD for 95% CI), and Mean ± 3 SD covers 99.7% of values.
- Standard Error of Mean (SEM = SD / √n) measures sampling variability; 95% Confidence Interval for sample mean is calculated as Mean ± 1.96 × SEM.
- WHO 30x7 Cluster Sampling is the recommended probability sampling technique for evaluating community immunization coverage.
- Unpaired t-test compares continuous means of 2 independent groups; Paired t-test compares continuous means before and after in the same subjects; Chi-square test compares proportions among categorical groups.
Biostatistics: Sampling, Central Tendency & Significance Testing
1. Classification of Variables / Data Types
Data collected in biomedical studies are broadly classified into two categories:
+---------------------------------------------------------------------------------------------------+
| TYPES OF BIOMEDICAL DATA |
+---------------------------------------------+-----------------------------------------------------+
| QUALITATIVE (CATEGORICAL) | QUANTITATIVE (NUMERICAL) |
+----------------------+----------------------+----------------------+------------------------------+
| Nominal | Ordinal | Discrete | Continuous |
+----------------------+----------------------+----------------------+------------------------------+
| Unordered categories | Ordered ranks/scales | Whole counts only | Measured on real continuum |
| • Blood Group (A/B/O)| • Cancer Stage (I-IV)| • Number of children | • Serum Bilirubin (mg/dL) |
| • Sex (Male/Female) | • Socioeconomic status| • Pulse rate/min | • Hemoglobin (g/dL) |
| • Eye color | • Severity (Mild-Mod)| • Bed occupancy count| • Height (cm), Weight (kg) |
+----------------------+----------------------+----------------------+------------------------------+
2. Measures of Central Tendency & Skewness
A. Summary Measures
- Mean (Arithmetic Average): Sum of all observations divided by total count ($n$). Highly sensitive to extreme values (outliers). Best measure of central tendency for normally distributed continuous data.
- Median (50th Percentile): The middle value when data are arranged in ascending order. Unaffected by extreme outliers. Best measure for skewed numerical data or ordinal data.
- Mode: The most frequently occurring value in a dataset. Useful for nominal data or identifying bimodal distributions.
B. Relationships in Skewed Distributions
Examples: Serum IgE levels and income distributions exhibit positive skew. Age at death in developed nations exhibits negative skew.
3. Measures of Dispersion (Variability)
- Range: Difference between maximum and minimum values. Highly unstable.
- Interquartile Range (IQR): $Q3 - Q1$ (75th percentile minus 25th percentile). Accompanies Median for skewed data.
- Standard Deviation (SD): Measures the scatter of individual observations around the sample mean:
- Standard Error of Mean (SEM): Measures the sampling variability of sample means around the true population mean:
- 95% Confidence Interval (95% CI): The range within which the true population parameter lies with 95% probability:
- Coefficient of Variation (CV): Relative measure of dispersion expressed as a percentage, enabling comparison between variables measured in different units (e.g., comparing height in cm vs weight in kg):
4. Normal (Gaussian) Distribution
The normal distribution is a symmetrical, bell-shaped continuous probability distribution defined by parameters $\mu$ (mean) and $\sigma$ (standard deviation).
+---------------------------------------------------------------------------------------------------+
| GAUSSIAN NORMAL DISTRIBUTION CURVE |
+---------------------------------------------------------------------------------------------------+
| | |
| / | \ |
| / | \ |
| / | \ |
| / | \ |
| / | \ |
| / | \ |
| | | | |
| ------------+---------+---------------+---------------+---------+------------ |
| μ-3σ μ-2σ μ-1σ μ μ+1σ μ+2σ μ+3σ |
| |<------------ 68.27% ------------>| |
| |<----------------------- 95.45% ----------------------->| |
| |<--------------------------------- 99.73% -------------------------------->| |
+---------------------------------------------------------------------------------------------------+
- Area Coverage Limits:
- $\text{Mean} \pm 1 SD$ encompasses 68.27% of values.
- $\text{Mean} \pm 2 SD$ encompasses 95.45% of values (specifically $\pm 1.96 SD$ encompasses 95.0%).
- $\text{Mean} \pm 3 SD$ encompasses 99.73% of values.
5. Sampling Techniques
A. Probability Sampling (Every element has known non-zero selection chance)
- Simple Random Sampling: Every individual in the sampling frame has an equal chance of selection (lottery or random number generator).
- Systematic Random Sampling: Every $k^{\text{th}}$ unit is selected after a random starting point ($k = N/n$).
- Stratified Random Sampling: Population is divided into homogeneous strata (e.g., age groups, socioeconomic strata), and a random sample is drawn from each stratum. Reduces sampling error.
- Cluster Sampling: Population is divided into heterogeneous clusters (e.g., villages, schools). A random selection of clusters is chosen, and all individuals within selected clusters are sampled.
- WHO 30x7 Cluster Sampling: Used globally for immunization coverage surveys. Evaluates 30 clusters of 7 children each aged 12-23 months (total $N = 210$).
- Multistage Sampling: Sampling is carried out in successive hierarchical stages (e.g., State $\rightarrow$ District $\rightarrow$ Village $\rightarrow$ Household).
B. Non-Probability Sampling
Includes Convenience, Purposive/Judgemental, Quota, and Snowball Sampling (used for hard-to-reach populations like IV drug users or sex workers).
6. Hypothesis Testing & Choice of Statistical Tests
Error Matrix in Hypothesis Testing
| True State of Nature \ Decision | Fail to Reject Null Hypothesis ($H_0$) | Reject Null Hypothesis ($H_0$) |
|---|---|---|
| Null Hypothesis ($H_0$) is True | Correct Decision ($1 - \alpha$) | Type I Error ($\alpha$) (False Positive) |
| Null Hypothesis ($H_0$) is False | Type II Error ($\beta$) (False Negative) | Correct Decision (Power $1 - \beta$) |
- p-value: The probability of obtaining the observed result (or more extreme) purely by chance if $H_0$ is true. If $p < 0.05$, $H_0$ is rejected.
Statistical Test Selection Algorithm
| Variable Type & Comparison Goal | Number of Groups / Samples | Parametric Test (Normal Data) | Non-Parametric Test (Skewed Data) |
|---|---|---|---|
| Compare Means (Continuous Data) | 2 Independent Groups | Unpaired Student's t-test | Mann-Whitney U test |
| Compare Means (Continuous Data) | 2 Paired Samples (Before & After) | Paired Student's t-test | Wilcoxon signed-rank test |
| Compare Means (Continuous Data) | $> 2$ Independent Groups | One-way ANOVA (F-test) | Kruskal-Wallis test |
| Compare Proportions (Categorical Data) | $\ge 2$ Independent Groups | Chi-Square Test ($\chi^2$) | Fisher's Exact Test (Cell $E < 5$) |
In a study of 100 healthy young adults, mean systolic blood pressure was 120 mmHg with a Standard Deviation of 10 mmHg. Assuming a normal Gaussian distribution, approximately how many individuals have a systolic BP between 100 mmHg and 140 mmHg?
Which of the following probability sampling methods is officially recommended by the World Health Organization (WHO) for assessing community immunization coverage?
A clinical investigator measures serum cholesterol levels in 50 hyperlipidemic patients before starting Statin therapy and again after 12 weeks of treatment. To test whether the reduction in mean cholesterol is statistically significant, which statistical test is most appropriate?
Rejecting a true Null Hypothesis (H₀) when there is in fact no real difference or effect in the population is termed: