10.8 Literature Evaluation: Research Methods, Statistics & Stability Evidence
Key Takeaways
Tertiary compatibility and stability references are starting points; for BUD decisions, check the primary study's drug, concentration, diluent, container, temperature and method against the actual CSP.
A stability study can support a BUD only if it uses a validated stability-indicating method, meaning one shown by forced degradation to separate drug from degradants, and tests the same formulation and container materials under the same storage conditions.
Chemical stability is commonly defined as keeping at least 90% of the initial concentration. For first-order degradation, t90 = 0.105 ÷ k.
Physical compatibility studies (visual, Tyndall beam, turbidity, particle counts) apply only to the concentrations and conditions tested. 'Compatible' in one study does not carry over to other conditions.
Clinical studies are judged by design, bias, statistical versus clinical significance, and absolute effect: with an absolute risk reduction (ARR) of 2 percentage points, the number needed to treat (NNT) is 1 ÷ 0.02 = 50.
10.8 Literature Evaluation: Research Methods, Statistics & Stability Evidence
BUDs, concentrations, compatibility decisions and new services all rest on published evidence. The BCSCP outline lists literature evaluation as practice-management task 3B2, covering research methods, statistics and whether a study applies to your patients and preparations. For a compounding specialist, the most common literature is the stability or compatibility study, and USP <797> requires Category 3 BUDs to rest on stability-indicating evidence.
Levels of information
| Level | Examples | How to use it |
|---|---|---|
| Tertiary | Trissel's Handbook on Injectable Drugs, King Guide to Parenteral Admixtures, ASHP's Extended Stability for Parenteral Drugs, drug-information databases, Stabilis | Quick screening; check the entry's cited primary study before assigning a BUD |
| Secondary | PubMed, Embase, International Pharmaceutical Abstracts | Find primary studies |
| Primary | Stability and compatibility studies, clinical trials, cohort studies | The actual evidence; judge its methods |
USP also publishes a Stability Study Reference Document that explains when a stability study is suitable for assigning Category 3 BUDs.
Judging a stability study for BUD use
Checklist, with every answer needing to be yes:
- Stability-indicating method. Is it HPLC or similar, shown by forced degradation (acid, base, oxidation, heat, light) to separate the drug from its degradants? A simple UV absorbance method may read degradants as intact drug.
- Method validation. Were specificity, accuracy, precision, linearity and range shown, for example against USP <1225>?
- Same formulation. Same drug (and salt form), concentration range, diluent, excipients and pH.
- Same container materials. PVC, polyolefin, glass and syringe type all matter, because of sorption and leaching.
- Same storage conditions. Temperature, light exposure, and any shift from frozen to refrigerated to room temperature.
- Adequate time points and replicates. The time points must cover the BUD being assigned, with at least 3 replicates.
- Physical stability too. pH, color, clarity, particulate matter and precipitation.
- Acceptance criteria. Typically at least 90% of the initial concentration remaining, and no clinically significant degradants. Stricter criteria apply if a degradant is toxic.
Remember that stability studies usually say nothing about sterility. The BUD must also respect the USP <797> category limits and sterility testing requirements. For Category 3, USP's FAQ allows a published or unpublished study, but only when the formulation, procedure and container are exactly the same as the CSP's.
Worked kinetics: t90 for first-order degradation
For first-order degradation, . The time until 90% remains is
If a study reports at 5 °C, then days. Better studies estimate shelf life from the lower 95% confidence limit of the regression line, not just the mean, which gives a more conservative date.
Judging a compatibility study
- Physical compatibility is judged visually (against a light and a dark background, and with a Tyndall beam), by turbidity measurement, and by particle counting. A lack of visible change does not rule out microprecipitates.
- Chemical compatibility needs a stability-indicating assay of each drug in the mixture.
- Results apply only to the concentrations, ratios, diluents and contact times tested. "Compatible at Y-site for 4 hours at 1 mg/mL" says nothing about 5 mg/mL or about 24 hours in the same bag.
- Conflicting reports are common. Weigh the method quality, not just the number of studies.
Clinical study designs
| Design | Strength | Example question in sterile therapy |
|---|---|---|
| Systematic review and meta-analysis | Pools the best evidence | Do CHG-alcohol dressings reduce CLABSI? |
| Randomized controlled trial | Tests cause and effect; limits confounding | Extended versus intermittent beta-lactam infusion |
| Cohort study | Follows exposure to outcome; good for harms | Outcomes of home PN patients by catheter type |
| Case-control study | Efficient for rare outcomes | Risk factors in a fungal outbreak |
| Case series or report | Signal detection | New incompatibility or adverse event |
Also check for bias (selection, performance, detection, attrition), confounding, blinding, sample size and power, and who funded the study.
Core statistics
- Mean, standard deviation and coefficient of variation describe precision, for example the ACD accuracy spread.
- Confidence interval (CI): a 95% CI for a difference that crosses 0 (or a ratio CI that crosses 1) means the result is not statistically significant.
- p-value: the probability of results this extreme if there were truly no effect. is conventional. Statistical significance is not the same as clinical importance.
- Type I error is a false positive (alpha); type II error is a false negative (beta). Power is 1 − beta, usually targeted at 80% or more.
- Relative versus absolute risk: if CLABSI falls from 4% to 2%, the relative risk reduction is 50%, but the absolute risk reduction (ARR) is 2 percentage points, and .
- Sensitivity and specificity matter when judging tests, such as rapid microbiological methods for sterility.
Applicability: does this study fit my CSP and my patients?
Ask whether the population, setting, product, container and conditions match yours. Separate internal validity (was the study done well?) from external validity (does it apply here?). A well-done neonatal study in glass syringes does not automatically support an adult elastomeric pump BUD.
A stability study reports first-order degradation of a drug in 0.9% sodium chloride in polyolefin bags at 5 °C, with a rate constant k = 0.0035 per day. What is the estimated time until 90% of the initial concentration remains (t90)?
About 10 days
About 20 days
About 30 days
About 90 days
Which study best supports a 60-day refrigerated Category 3 BUD for a drug at 2 mg/mL in 0.9% sodium chloride in polyolefin bags?
A study at 2 mg/mL in 0.9% sodium chloride in polyolefin bags at 2-8 °C, using a forced-degradation-validated stability-indicating HPLC method, with triplicate samples through 60 days showing at least 90% remaining and no physical change
A UV absorbance study of the same drug at 10 mg/mL in glass bottles at room temperature
A drug-information database entry listing the drug as 'stable 60 days' with no citation
A study showing the drug looks visually compatible for 4 hours at a Y-site
A trial finds that a new catheter-care bundle lowers the CLABSI rate from 4% to 2%. What are the absolute risk reduction and the number needed to treat?
ARR 50%, NNT 2
ARR 2 percentage points, NNT 50
ARR 4 percentage points, NNT 25
ARR 2 percentage points, NNT 2
Sections you finish are checked off in the contents.