13.1 Descriptive Statistics, Distributions, and Parametric vs Non-Parametric Tests
Key Takeaways
Categorical data comprises nominal (unordered classifications such as ABO blood group) and ordinal variables (ranked hierarchy without constant mathematical intervals such as ASA physical status or Mallampati grade); numerical data comprises discrete counts and continuous scales (interval vs ratio).
In a symmetrical Gaussian (normal) distribution, mean = median = mode; data dispersion is described by standard deviation (), where encompasses , encompasses , and encompasses of observations.
For skewed distributions, central tendency is reported using the median and dispersion using the interquartile range (, the middle of data); reporting the mean and for skewed variables severely distorts clinical interpretation.
The Standard Error of the Mean () quantifies the precision of a sample mean estimating the true population mean; it must never be substituted for in clinical baseline tables to deceptively minimize apparent biological variance.
Parametric tests (unpaired/paired Student's t-test, ANOVA) require continuous, normally distributed data with equal variance (homoscedasticity); if distribution assumptions are violated, non-parametric equivalents based on rank sums (Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis) must be employed.
13.1 Descriptive Statistics, Distributions, and Parametric vs Non-Parametric Tests
Biostatistics forms the quantitative bedrock of modern evidence-based anaesthesia and intensive care medicine. Clinicians must possess a rigorous command of data classification, distribution geometries, summary statistics, and hypothesis testing to critically evaluate scientific literature, interpret clinical monitoring data, and avoid common methodological pitfalls.
1. Taxonomy of Biological and Clinical Data
Clinical data obtained in anaesthetic practice are broadly categorized into qualitative (categorical) and quantitative (numerical) variables. Selecting the correct descriptive statistics and inferential tests depends entirely on this initial classification.
[ CLINICAL DATA ]
|
+----------------------------+----------------------------+
| |
[ Qualitative / Categorical ] [ Quantitative / Numerical ]
| |
+----+----+ +----+----+
| | | |
[ Nominal ] [ Ordinal ] [ Discrete ] [ Continuous ]
(Unordered) (Ranked) (Counts) |
+----+----+
| |
[ Interval ] [ Ratio ]
(No true 0) (True absolute 0)
Qualitative (Categorical) Data
- Nominal Data: Mutually exclusive categories that have no inherent mathematical order or ranking. Arithmetic operations cannot be performed.
- Examples: ABO blood group (A, B, AB, O), biological sex, eye colour, surgical specialty or anatomical surgical site.
- Binary / Dichotomous Data: A specific subtype with exactly two categories (e.g., survival vs death, post-dural puncture headache present vs absent).
- Ordinal Data: Categorical data organized into an ordered or ranked hierarchy. Although the sequence is ordered, the mathematical intervals between consecutive ranks are neither equal nor quantifiable.
- Examples: ASA physical status classification (I to VI), Mallampati airway classification (Class I to IV), New York Heart Association (NYHA) functional class (I to IV), Cormack-Lehane laryngoscopy grade (1 to 4), Glasgow Coma Scale (GCS).
- Crucial Rule: Because intervals are non-linear, calculating an arithmetic mean of ordinal data (e.g., "an average Mallampati score of 2.4") is mathematically invalid. Ordinal variables must be summarized using medians with interquartile ranges or proportions.
Quantitative (Numerical) Data
- Discrete Data: Integer values representing countable entities with no intermediate fractional values.
- Examples: Number of intubation attempts, number of central line puncture passes, number of episodes of postoperative vomiting, number of days alive and free of mechanical ventilation.
- Continuous Data: Measurements along an unbroken numerical continuum where infinite fractional values are theoretically possible, constrained only by measurement precision.
- Interval Scale: Continuous data with equal intervals between scale units but no true, non-arbitrary absolute zero. Ratios between numbers are meaningless.
- Example: Temperature measured in degrees Celsius (°C) or Fahrenheit (°F). is arbitrarily defined as the freezing point of water, not the total absence of thermodynamic heat energy; is not "twice as hot" as .
- Ratio Scale: Continuous data featuring both equal intervals and a true physical absolute zero representing the total absence of the measured entity. Mathematical multiplication and division are fully valid.
- Examples: Mean arterial blood pressure (mmHg), cardiac output (L/min), patient weight (kg), alveolar minimum alveolar concentration (MAC), serum potassium (mmol/L), temperature in Kelvin (K). A cardiac output of is legitimately twice as large as .
- Interval Scale: Continuous data with equal intervals between scale units but no true, non-arbitrary absolute zero. Ratios between numbers are meaningless.
| Data Type | Primary Defining Feature | Central Tendency | Dispersion Metric | Typical Anaesthetic Examples |
|---|---|---|---|---|
| Nominal | Unordered categories | Mode | Frequency / Proportions | Blood group, surgical specialty |
| Ordinal | Ordered ranks, unequal intervals | Median | Interquartile Range () | ASA status, Mallampati grade, NYHA class |
| Discrete | Countable integers | Median or Mean | or Range | Intubation attempts, defibrillation shocks |
| Continuous (Interval) | Equal units, arbitrary zero | Mean (if Gaussian) | Standard Deviation () | Temperature in °C or °F |
| Continuous (Ratio) | Equal units, true absolute zero | Mean (if Gaussian) | Standard Deviation () | Blood pressure, cardiac output, MAC |
2. Measures of Central Tendency and Data Dispersion
Descriptive statistics summarize data distributions using a measure of central location (location of the midpoint) paired with an appropriate measure of statistical dispersion (spread).
Measures of Central Tendency
- Arithmetic Mean (): The sum of all observations divided by the sample size ():
- Properties: Incorporates every datum; mathematically tractable for algebraic derivations. Highly vulnerable to distortion by extreme outliers and skewed distributions.
- Median: The middle value when all observations are arranged in ascending numerical order (). If is even, it is the arithmetic average of the two middle values.
- Properties: Robust against extreme outliers and severe skewness. Preferred summary statistic for non-normally distributed continuous data and ordinal scales.
- Mode: The most frequently occurring value in the sample.
- Properties: Sole measure of central tendency applicable to nominal data. A distribution may be unimodal, bimodal (two distinct peaks), or multimodal.
Measures of Dispersion
- Range: The arithmetic difference between the maximum and minimum observations (). Highly unstable because it relies entirely on the two most extreme values.
- Interquartile Range (): The difference between the 75th percentile (third quartile, ) and the 25th percentile (first quartile, ):
- The encompasses the central of all observations. It is completely unaffected by extreme outliers in the outer tails and is paired standardly with the median.
- Variance ( for population, for sample): The average squared deviation of individual values from the arithmetic mean:
- The denominator uses Bessel's correction () to provide an unbiased estimate of the true population variance.
- Standard Deviation ( or ): The positive square root of the variance ():
- Returns the dispersion metric to the original units of measurement. Correctly paired with the mean only when data follow a Gaussian distribution.
Box-and-Whisker Plots (Tukey Boxplots)
Box-and-whisker plots display five-number summaries:
- Horizontal line inside the box: Median ().
- Lower and upper box hinges: 25th percentile () and 75th percentile (); box height represents the .
- Whiskers: Extend to the furthest data point within from the hinges.
- Individual dots outside whiskers: Outliers exceeding from the quartiles.
3. The Gaussian (Normal) Distribution and the Empirical Rule
The normal or Gaussian distribution is the fundamental continuous probability distribution in biological statistics. It is generated when an outcome is influenced by the additive sum of numerous independent random variables (Central Limit Theorem).
Normal Curve
/|\
/ | \
/ | \
/ | \
/ | \
---+-----+-----+---
-2 Mean +2
Median
Mode
[<-------------- 68.3% (±1 SD) -------------->]
[<----------------------- 95.0% (±1.96 SD) ----------------------->]
[<----------------------------- 99.7% (±3 SD) ----------------------------->]
Mathematical Characteristics
- Symmetrical, bell-shaped, asymptotic to the horizontal axis (never touches zero).
- Unimodal with inflection points located exactly at .
- Coincidence of central indices: .
- Skewness equals zero; kurtosis (peakedness) equals 3 (mesokurtic).
The Empirical (68-95-99.7) Rule
In any genuine Gaussian distribution with mean and standard deviation :
- encompasses of all observations.
- encompasses exactly of all observations (often approximated as ).
- encompasses exactly of all observations.
- encompasses of all observations.
The Standardized Normal Deviate (-Score)
Any value from a normal distribution can be transformed into a standard score (), expressing distance from the mean in standard deviation units ():
4. Skewed Distributions in Anaesthesia and Intensive Care
Biological data in acute care medicine frequently deviate from Gaussian symmetry. Skewness refers to asymmetry in the tail of the probability density function.
Positive (Right) Skew Negative (Left) Skew
/\ /\
/ \ / \
/ \ / \
/ \___ / \
_/ \___ ___/ \_
---+---+---+---------+--- ---+---------+---+---+---
Mode Med Mean Mean Med Mode
[Tail to the RIGHT] [Tail to the LEFT]
Positive (Right) Skew
- The right-hand tail (toward higher positive values) is prolonged and drawn out.
- Mathematical Order: . The arithmetic mean is pulled towards the extreme positive values in the tail.
- Classic Clinical Examples in Anaesthesia & ICU:
- Intensive care unit length of stay (days): Most patients are discharged within 2-4 days, but a small cohort with multi-organ failure remains for 60+ days.
- Postoperative patient-controlled analgesia (PCA) morphine consumption (mg).
- Post-anaesthesia care unit (PACU) discharge delay time.
- Plasma concentrations of troponin, ferritin, and C-reactive protein.
- Data Transformation: Right-skewed data can often be converted to a Gaussian distribution via logarithmic transformation (taking the natural logarithm ), allowing parametric testing.
Negative (Left) Skew
- The left-hand tail (toward lower values) is prolonged and drawn out.
- Mathematical Order: . The arithmetic mean is pulled downward by extreme low values.
- Classic Clinical Examples:
- Age at elective joint replacement surgery (e.g., total hip arthroplasty: clustered between 65-80 years, with a sparse tail extending down to young trauma/rheumatoid patients).
- Arterial oxygen saturation () in healthy post-surgical patients on room air (clustered between 96-99%, with an elongated tail of hypoxaemic outliers extending down to 80%).
5. Standard Deviation () vs Standard Error of the Mean ()
Confusing standard deviation with standard error is one of the most pervasive errors in biomedical reporting and a frequent target on primary examinations.
Standard Deviation ()
- Definition: Quantifies the dispersion or biological variability of individual measurements around the sample mean.
- Formula: .
- Sample Size Behavior: As sample size () increases, does not systematically decrease; it converges more accurately on the true, constant population standard deviation ().
- Clinical Purpose: Used strictly for descriptive statistics to communicate biological scatter among patients.
Standard Error of the Mean ()
- Definition: Quantifies the precision with which the sample mean () estimates the unknown true population mean (). It represents the standard deviation of the sampling distribution of means.
- Formula:
- Sample Size Behavior: Inversely proportional to the square root of . As sample size increases four-fold (), the is halved ().
- Clinical Purpose: Used strictly for inferential statistics to construct confidence intervals around parameter estimates.
Clinical Trap: The Illicit Use of
Because , is inevitably smaller than by a factor of . For example, in a study of patients with a mean arterial pressure of and , the is only .
Authors sometimes intentionally report baseline demographic tables as () instead of () because the narrow error bars deceptively make the clinical measurements appear tightly controlled. Reporting to describe sample variability is scientifically invalid. Baseline patient characteristics must always be summarized using (for Gaussian data) or (for skewed data).
6. Parametric versus Non-Parametric Hypothesis Tests
Hypothesis tests are classified based on the distribution assumptions required to ensure mathematical validity.
Parametric Test Assumptions
To utilize parametric tests without inflating Type I or Type II error rates, four criteria must be satisfied:
- Continuous Data: Measured on an interval or ratio scale.
- Normal Distribution: The underlying populations sampled must be normally distributed (evaluated using the Shapiro-Wilk or Kolmogorov-Smirnov test, or graphical Q-Q plots).
- Homoscedasticity: Variances across compared groups must be approximately equal (evaluated using Levene's test or Bartlett's test).
- Independence: Observations within and between groups must be independent.
Non-Parametric Tests (Distribution-Free)
When data are ordinal, discrete, or continuous with marked skewness or unequal variances, non-parametric tests must be selected. These methods do not evaluate raw numerical parameters; instead, they rank-order the observations and evaluate rank sums or median differences.
- Advantage: Highly robust against outliers and distribution skewness; valid for small sample sizes where normality cannot be confirmed.
- Disadvantage: Approximately less statistical power than parametric tests when data are genuinely Gaussian.
Test Selection Matrix
| Clinical Research Question | Parametric Test (Normal Continuous) | Non-Parametric Equivalent (Ranked / Skewed / Ordinal) |
|---|---|---|
| Compare 2 Independent Groups | Two-sample Unpaired Student's t-test | Mann-Whitney U test (Wilcoxon rank-sum test) |
| Compare 2 Paired / Matched Groups | Paired Student's t-test | Wilcoxon signed-rank test |
| Compare Independent Groups | One-way Analysis of Variance (ANOVA) | Kruskal-Wallis test |
| Compare Repeated Measures | Repeated-Measures ANOVA | Friedman test |
| Correlation between 2 Variables | Pearson product-moment () | Spearman rank correlation () |
Note on Post-Hoc Testing: When an ANOVA or Kruskal-Wallis test reveals a statistically significant omnibus difference across three or more groups, pairwise comparisons require post-hoc adjustments (e.g., Bonferroni correction, Tukey's honestly significant difference [HSD], or Dunn's test) to control the family-wise error rate.
7. Categorical Data Analysis: Contingency Tables
When evaluating frequencies and proportions of categorical variables across groups, contingency tables ( or ) are assessed.
Chi-Squared () Test of Independence
Evaluates whether an observed distribution of frequencies differs significantly from the theoretical distribution expected under the null hypothesis of independence.
- Formula: Where is the observed cell count and is the expected cell frequency:
- Degrees of Freedom (): For an table, .
- The Cochran Criteria (Assumptions):
- Observations must be independent (each subject contributes to only one cell).
- In tables with , no cell should have an expected frequency , and at least of cells must have .
- For a standard table, all four cells must have an expected frequency .
Fisher's Exact Test
When expected cell frequencies in a contingency table fall below 5 (), the chi-squared distribution's continuous approximation breaks down, producing unreliable P-values.
- Mechanism: Calculates the exact hypergeometric probability of observing the specific cell frequencies, holding marginal totals fixed.
- Clinical Application: Indicated for small sample sizes, rare clinical complications, or any table where any expected cell count is .
A clinical audit of 200 patients undergoing elective laparotomy evaluates ASA physical status classification (grades I to IV), baseline mean arterial pressure in mmHg, and post-anaesthesia care unit (PACU) length of stay in minutes. Which statement correctly identifies the data classification and optimal reporting metrics for these variables?
ASA status is ordinal data best reported as median (IQR) or proportions; PACU stay is right-skewed continuous data best summarized by median (IQR); mean arterial pressure is continuous ratio data reported as mean (SD) if normally distributed.
ASA status is nominal categorical data best reported as mean (SD); PACU stay is discrete count data requiring parametric paired t-test evaluation; mean arterial pressure is continuous interval data reported as mean (SEM).
ASA status is continuous ordinal data best described by mean and range; PACU stay is normally distributed interval data requiring mean (SD); mean arterial pressure is discrete ratio data summarized by the mode and range.
ASA status is nominal qualitative data best summarized by standard error; PACU stay is left-skewed continuous data reported as mean (SD); mean arterial pressure is ordinal data best reported as median (IQR).
In a trial of 100 hypertensive patients undergoing carotid endarterectomy, systolic blood pressure was measured before induction. The sample mean was 160 mmHg with a standard deviation (SD) of 20 mmHg. Which statement correctly distinguishes the Standard Deviation from the Standard Error of the Mean (SEM) in this study?
The SEM is 20 mmHg and describes the physiological dispersion of systolic blood pressure among individual patients in the cohort.
The SD (20 mmHg) describes variability between patients; the SEM (2.0 mmHg) describes how precisely the sample mean estimates the population mean.
The SEM is 2.0 mmHg and indicates that 95% of individual patient blood pressures in the clinic will fall between 156.08 and 163.92 mmHg.
As the study sample size increases from 100 to 400 patients, the SD will decrease four-fold to 5 mmHg, whereas the SEM will remain constant at 20 mmHg.
An investigator compares postoperative visual analogue scale (VAS) pain scores (0 to 100 mm) at 24 hours between two independent groups of 35 patients receiving either intravenous ketamine or placebo. The pain scores in both groups exhibit marked right-skewness and fail the Shapiro-Wilk normality test. Which statistical hypothesis test is the most appropriate to compare the primary outcome between the two groups?
Paired Student's t-test on log-transformed scores
One-way Analysis of Variance (ANOVA)
Mann-Whitney U test (Wilcoxon rank-sum)
Chi-squared () test of independence
Sections you finish are checked off in the contents.