2.5 Data Collection, Presentation, Visualization & Statistical Analysis
Key Takeaways
- Statistical Process Control (SPC) control charts distinguish common cause random variation from special cause assignable variation.
- p-charts monitor proportional data (percentages), whereas u-charts track rate data per unit of exposure (e.g., HAIs per 1,000 device days).
- A p-value < 0.05 indicates statistical significance, rejecting the null hypothesis of no difference.
- If a 95% Confidence Interval for Relative Risk or Odds Ratio includes 1.0, the association is not statistically significant.
- In skewed distributions, the median provides a more valid measure of central tendency than the arithmetic mean.
2.5 Data Collection, Presentation, Visualization & Statistical Analysis
Transforming raw surveillance counts into actionable clinical knowledge requires mastery of data visualization tools and statistical analysis. Infection Preventionists must present data clearly to multidisciplinary quality committees, distinguish random process variation from true clinical degradation, and apply inferential statistics to determine whether observed changes in infection rates are statistically meaningful.
1. Data Presentation & Visualization Tools
Selecting the correct chart type depends on the underlying measurement scale and analytical purpose:
Run Charts vs. Control Charts
- Run Chart: A line graph displaying continuous data points plotted over time (e.g., monthly infection counts over 12 months). A horizontal center line representing the median or mean is included. Run charts display temporal trends, shifts, and cycles without formal statistical control limits.
- Control Chart (Statistical Process Control / SPC): A run chart enhanced with mathematically calculated upper and lower statistical boundaries:
- Center Line (CL): Represents the historical process mean or median.
- Upper Control Limit (UCL): Positioned $3 \text{ standard deviations } (+3\sigma)$ above the center line.
- Lower Control Limit (LCL): Positioned $3 \text{ standard deviations } (-3\sigma)$ below the center line.
UCL -------------------------------------------------- (+3 Sigma)
CL ────────────────────────────────────────────────── (Process Mean)
LCL -------------------------------------------------- (-3 Sigma)
Types of Control Charts in Infection Prevention
- p-Chart (Proportion Chart): Used for attribute data expressed as proportions or percentages (e.g., monthly hand hygiene compliance rates, proportion of surgical patients receiving timely prophylactic antibiotics).
- u-Chart (Rate Chart): Used for rate data where the denominator varies (e.g., monthly CLABSI rates per $1,000$ central line-days).
- c-Chart (Count Chart): Used for absolute defect counts when the sample size/denominator remains constant (e.g., monthly needle-stick injury counts).
Interpreting SPC Charts: Variation Types
- Common Cause Variation: Inherent, natural background noise expected within a stable healthcare process. Points fluctuate randomly between the UCL and LCL without distinct patterns. Requires systemic process redesign to improve.
- Special Cause Variation: Non-random, assignable variation caused by an external factor or process failure. Indicated by SPC rule breaches:
- Any single data point falling outside the UCL or LCL boundaries.
- Shift: 8 consecutive data points falling entirely above or below the center line.
- Trend: 6 consecutive data points continuously increasing or continuously decreasing.
Bar Graphs vs. Histograms
- Bar Graph: Displays categorical/discrete variables (e.g., infection counts by clinical specialty: ICU, Stepdown, Med-Surg). Distinct spaces exist between bars to reflect discrete, non-continuous categories.
- Histogram: Displays continuous quantitative numerical variables (e.g., patient age distribution, duration of mechanical ventilation). Bars touch with no spaces, illustrating continuous frequency distributions.
2. Descriptive Statistics: Central Tendency & Dispersion
Descriptive statistics summarize key characteristics of a dataset:
Measures of Central Tendency
- Mean (Arithmetic Average): Sum of all values divided by total count ($N$). Sensitive to extreme outliers; best suited for symmetric, normally distributed data.
- Median (Middle Value): The 50th percentile value when data are ranked sequentially. Robust against extreme outliers; preferred for skewed distributions (e.g., post-operative length of stay).
- Mode: The most frequently occurring value in a dataset.
Distribution Skewness Pearl:
- Symmetrical (Normal) Distribution: $\text{Mean} = \text{Median} = \text{Mode}$.
- Positively (Right) Skewed: $\text{Mean} > \text{Median} > \text{Mode}$ (tail extends right toward higher values).
- Negatively (Left) Skewed: $\text{Mean} < \text{Median} < \text{Mode}$ (tail extends left toward lower values).
Measures of Dispersion
- Standard Deviation (SD / $\sigma$): Quantifies the spread of data points around the mean. In a normal distribution:
- $\pm 1\text{ SD}$ encompasses $68.27%$ of data points.
- $\pm 2\text{ SD}$ encompasses $95.45%$ of data points.
- $\pm 3\text{ SD}$ encompasses $99.73%$ of data points (foundation of SPC control limits).
3. Inferential Statistics & Significance Testing
Inferential statistics allow IPs to draw conclusions about a broader population based on sample data.
The Hypothesis Testing Framework
- Null Hypothesis ($H_0$): Assumes no true difference, association, or effect exists between groups (any observed difference is due purely to random chance).
- Alternative Hypothesis ($H_1$): Assumes a true, non-random difference or association exists.
- p-Value: The probability of obtaining a sample result as extreme as (or more extreme than) the observed data, assuming the null hypothesis is true. By convention, $\alpha = 0.05$.
- $p < 0.05$: Reject $H_0$. Finding is statistically significant.
- $p \ge 0.05$: Fail to reject $H_0$. Finding is not statistically significant.
95% Confidence Intervals (95% CI)
A Confidence Interval provides a range of plausible values for the true population parameter with a specified level of certainty ($95%$).
Critical CBIC Rules for Interpreting 95% CIs:
- Ratio Metrics (Relative Risk, Odds Ratio, SIR):
- If the 95% CI includes 1.00 (e.g., $\text{RR} = 1.80$, $95% \text{ CI: } 0.90 - 3.40$), the result is NOT statistically significant ($p \ge 0.05$).
- If the 95% CI excludes 1.00 (e.g., $\text{RR} = 2.40$, $95% \text{ CI: } 1.30 - 4.20$), the result IS statistically significant ($p < 0.05$).
- Difference Metrics (Difference between two rates or means):
- If the 95% CI includes 0.00, the difference is NOT statistically significant.
Statistical Tests Matrix
| Statistical Test | Variable Types | Primary Application |
|---|---|---|
| Chi-Square ($\chi^2$) Test | Two categorical variables | Compares proportions between two independent groups (sample sizes $> 5$ per cell). |
| Fisher's Exact Test | Two categorical variables | Preferred over Chi-Square when cell counts are small (expected frequency $< 5$). |
| Student's t-Test | Continuous variable vs. 2 categorical groups | Compares means of two groups (e.g., length of stay between infected vs. uninfected). |
| ANOVA (Analysis of Variance) | Continuous variable vs. $\ge 3$ categorical groups | Compares continuous means across three or more groups. |
An Infection Preventionist constructs a p-chart to monitor monthly hand hygiene compliance rates across an acute care facility over 24 months. The process mean is 85%. In month 18, compliance drops to 62%, falling below the Lower Control Limit (LCL). How should this data point be interpreted?
In a case-control study investigating risk factors for surgical site infection following colon surgery, the calculated Odds Ratio for perioperative hypothermia is 2.80, with a 95% Confidence Interval of 0.85 to 5.40. Which conclusion is correct?
An Infection Preventionist is comparing the proportion of catheter-associated urinary tract infections between two intensive care units. In one unit, the cell counts in the 2x2 contingency table contain expected frequencies of less than 5. Which statistical test should be used to evaluate significance?
A dataset recording post-operative hospital length of stay (in days) for 100 surgical patients is heavily skewed to the right (positively skewed) by a few patients who experienced severe complications and stayed over 90 days. Which measure of central tendency provides the most accurate summary of typical length of stay?