16.1 Oncology Clinical Trial Designs, Endpoints & Statistical Power

Key Takeaways

  • Oncology Phase I trials utilize rule-based algorithms (classical 3+3, accelerated titration) or Bayesian model-based designs (Continuous Reassessment Method [CRM], BOIN) to characterize dose-limiting toxicities (DLTs) and define the maximum tolerated dose (MTD) or optimal biologic dose (OBD).
  • Simon's two-stage Phase II design minimizes expected sample size under the null hypothesis by permitting early termination for futility if fewer than r1 responses occur in the initial n1 evaluable patient cohort before proceeding to stage two.
  • Master protocols accelerate precision oncology: basket trials evaluate a single targeted therapy across multiple tumor histologies harboring a common genomic alteration, umbrella trials evaluate multiple biomarker-directed therapies within a single tumor type, and platform trials evaluate multiple therapies dynamically against a shared control.
  • Overall survival (OS) is the gold standard oncology endpoint reflecting unambiguous clinical benefit unconfounded by assessment bias, whereas progression-free survival (PFS) and overall response rate (ORR) serve as regulatory surrogate endpoints susceptible to investigator measurement bias and post-progression crossover confounding.
  • Statistical power in time-to-event oncology trials is strictly driven by the total number of observed events rather than total patient sample size alone; non-inferiority trials require pre-specifying a clinically justified non-inferiority margin (Δ) that preserves active comparator efficacy.
Last updated: August 2026

16.1 Oncology Clinical Trial Designs, Endpoints & Statistical Power

Clinical trial methodology and biostatistical evaluation are foundational competencies for Board-Certified Oncology Pharmacists (BCOPs). The oncology therapeutic landscape evolves rapidly through innovative clinical trial designs, precision biomarker stratification, master protocols, and surrogate endpoint validation. Evaluating new antineoplastic therapies requires understanding study architectures, dose-finding algorithms, endpoint mechanics, error control, and sample size determinants.


1. Classical Clinical Trial Phases in Oncology

+---------------------------------------------------------------------------------------------------+
|                         CLASSICAL ONCOLOGY CLINICAL TRIAL PHASES                                  |
|                                                                                                   |
|   [PHASE 0] (Exploratory IND / Microdosing):                                                      |
|   - Sample: 10–15 patients. Subtherapeutic microdoses (1/100th of pharmacological dose).          |
|   - Objective: Early human PK/PD, target engagement, PET biodistribution. Zero therapeutic intent.|
|                                                                                                   |
|   [PHASE I] (Dose-Escalation & Safety / DLT Window):                                             |
|   - Sample: 20–80 patients (advanced/refractory solid tumors or hematologic malignancies).         |
|   - Objective: Determine Maximum Tolerated Dose (MTD), Dose-Limiting Toxicities (DLTs),           |
|     Pharmacokinetics (PK)/Pharmacodynamics (PD), and Recommended Phase 2 Dose (RP2D).            |
|                                                                                                   |
|   [PHASE II] (Therapeutic Exploratory / Efficacy Screening):                                      |
|   - Sample: 40–150 patients with specific tumor histology and molecular biomarker.                |
|   - Objective: Assess preliminary antitumor activity (ORR, DCR), confirm safety, evaluate PFS.   |
|   - Standard Architectures: Simon's Two-Stage Design, Single-Arm or Randomized Phase II.          |
|                                                                                                   |
|   [PHASE III] (Confirmatory Therapeutic / Comparative Efficacy):                                  |
|   - Sample: 300–3,000+ patients across multicenter international sites.                           |
|   - Objective: Demonstrate superiority or non-inferiority versus standard-of-care (SOC) on hard   |
|     clinical endpoints (Overall Survival [OS], Progression-Free Survival [PFS]).                  |
|   - Design: Randomized Controlled Trial (RCT), stratified, double-blind or open-label with BICR.  |
|                                                                                                   |
|   [PHASE IV] (Post-Marketing Surveillance & Real-World Evidence):                                 |
|   - Sample: Thousands of unselected real-world patients.                                          |
|   - Objective: Detect rare adverse events, long-term safety, off-label usage patterns, REMS.     |
+---------------------------------------------------------------------------------------------------+

Phase I Dose-Escalation Frameworks

Phase I oncology trials evaluate novel compounds in humans to establish the recommended Phase II dose (RP2D). Unlike non-oncology Phase I trials conducted in healthy volunteers, oncology Phase I trials enroll patients with advanced, refractory malignancies who have exhausted standard therapy options.

The Classical "3+3" Rule-Based Design

The standard rule-based Phase I design enrolls cohorts of 3 to 6 patients at predefined, escalating dose levels (frequently using a modified Fibonacci sequence where dose increments diminish: +100%, +67%, +50%, +40%, +33%):

+---------------------------------------------------------------------------------------------------+
|                         THE CLASSICAL "3+3" DOSE-ESCALATION ALGORITHM                             |
|                                                                                                   |
|   Enter 3 patients at Dose Level K                                                                |
|   |                                                                                               |
|   +---> 0 / 3 develop DLT =====> ESCALATE to Dose Level K+1 (Enroll 3 new patients)               |
|   |                                                                                               |
|   +---> 1 / 3 develops DLT ====> EXPAND Cohort K: Enroll 3 additional patients (Total = 6)       |
|   |     |                                                                                         |
|   |     +---> 1 / 6 develops DLT =====> ESCALATE to Dose Level K+1                                |
|   |     |                                                                                         |
|   |     +---> >=2 / 6 develop DLT ====> MTD EXCEEDED. Stop escalation.                            |
|   |                                     Dose Level K-1 is the Maximum Tolerated Dose (MTD).       |
|   |                                                                                               |
|   +---> >=2 / 3 develop DLT ====> MTD EXCEEDED. Stop escalation.                                  |
|                                   Dose Level K-1 is the Maximum Tolerated Dose (MTD).             |
+---------------------------------------------------------------------------------------------------+
  • Dose-Limiting Toxicity (DLT): Severe, treatment-related toxicities occurring during a protocol-defined observation window (typically Cycle 1, Days 1–21 or Days 1–28). Examples: Grade 4 neutropenia lasting >5–7 days, Grade 3/4 febrile neutropenia, Grade 4 thrombocytopenia or Grade 3 thrombocytopenia with bleeding, or any Grade $\ge 3$ non-hematologic organ toxicity (excluding nausea/vomiting/diarrhea responsive to optimal medical therapy).
  • Maximum Tolerated Dose (MTD): The highest dose level at which $\le 1$ of 6 patients ($\le 16.7%$) experiences a DLT.
  • Limitations of 3+3 Design: Treats many patients at subtherapeutic starting doses, slow escalation pace, poor estimation of true MTD precision, and assumes toxicity increases monotonically with dose (which may not apply to targeted TKIs or monoclonal antibodies).

Model-Based & Adaptive Phase I Designs

  • Continuous Reassessment Method (CRM): A Bayesian model-based design that dynamically updates the dose-toxicity probability curve after each patient outcome, assigning subsequent patients to the dose closest to the target toxicity probability (e.g., target DLT rate of 25–30%).
  • Bayesian Optimal Interval (BOIN) Design: Pre-calculates transparent dose-escalation, de-escalation, and retention boundaries based on observed DLT rates within the current cohort, combining the simplicity of rule-based designs with the statistical accuracy of Bayesian models.
  • Maximum Tolerated Dose (MTD) vs Optimal Biologic Dose (OBD): Targeted therapies and immunotherapies often achieve maximum target receptor saturation or immunologic activation at doses well below the MTD without causing dose-limiting toxicities. In these settings, trials establish the Optimal Biologic Dose (OBD) using pharmacodynamic biomarkers (e.g., target phosphorylation inhibition, immune cell activation) rather than escalating to maximum toxicity.

2. Phase II Trial Architectures: Simon's Two-Stage Design

Phase II trials determine whether an investigational antineoplastic exhibits sufficient antitumor activity to warrant Phase III evaluation. Because exposing patients to ineffective cytotoxic agents is unethical, Simon's Two-Stage Design is the industry standard for single-arm Phase II trials with binary endpoints (e.g., Objective Response Rate [ORR]).

+---------------------------------------------------------------------------------------------------+
|                         SIMON'S TWO-STAGE PHASE II CLINICAL TRIAL DESIGN                          |
|                                                                                                   |
|   HYPOTHESES:                                                                                     |
|   - Null Hypothesis (H0): Response Rate <= p0 (Unacceptable / Ineffective; e.g., p0 = 15%)         |
|   - Alternative Hypothesis (H1): Response Rate >= p1 (Desirable / Promising; e.g., p1 = 35%)      |
|   - Error Constraints: Type I error (alpha, false positive) and Type II error (beta, false neg)   |
|                                                                                                   |
|   [STAGE 1]:                                                                                      |
|   - Enroll Stage 1 cohort of n1 patients.                                                         |
|   - Evaluate confirmed objective responses (CR + PR):                                             |
|     * If responses <= r1 ======> TERMINATE TRIAL EARLY FOR FUTILITY (Reject H1).                  |
|     * If responses > r1  ======> PROCEED TO STAGE 2.                                              |
|                                                                                                   |
|   [STAGE 2]:                                                                                      |
|   - Enroll additional n2 patients (Total sample size N = n1 + n2).                                |
|   - Evaluate confirmed objective responses across all N patients:                                 |
|     * If total responses <= r ======> DRUG IS INEFFECTIVE (Fail to reject H0).                    |
|     * If total responses > r  ======> DRUG IS PROMISING (Reject H0; proceed to Phase III).        |
|                                                                                                   |
|   DESIGN VARIANTS:                                                                                |
|   - **Simon's Optimal Design:** Minimizes the Expected Sample Size (EN) under the null hypothesis |
|     (stops earliest if the drug is truly inactive).                                               |
|   - **Simon's Minimax Design:** Minimizes the Maximum Total Sample Size (N = n1 + n2)             |
|     (preferred when patient accrual is slow or rare disease).                                     |
+---------------------------------------------------------------------------------------------------+

3. Master Protocols & Precision Oncology Trial Architectures

Traditional clinical trials evaluate a single drug in a single disease histology. Precision oncology employs Master Protocols—coordinated clinical trial infrastructures that evaluate multiple investigational treatments, multiple patient subpopulations, or multiple disease types under a single overarching protocol framework.

+---------------------------------------------------------------------------------------------------+
|                    MASTER PROTOCOL TRIAL ARCHITECTURES IN PRECISION ONCOLOGY                      |
|                                                                                                   |
|   [1. BASKET TRIAL] (Histology-Agnostic / Single Biomarker -> Multiple Tumor Types)               |
|   - Concept: Evaluates ONE targeted therapy across MULTIPLE distinct tumor histologies            |
|     harboring the SAME specific molecular alteration.                                             |
|   - Examples:                                                                                     |
|     * Larotrectinib / Entrectinib for NTRK gene fusions across CRC, thyroid, sarcoma, NSCLC.      |
|     * Dabrafenib + Trametinib for BRAF V600E-mutated solid tumors (biliary, glioma, anaplastic).   |
|     * Pembrolizumab for MSI-High / dMMR or TMB-High solid tumors.                                 |
|                                                                                                   |
|   [2. UMBRELLA TRIAL] (Single Histology -> Multiple Biomarkers & Targeted Therapies)             |
|   - Concept: Evaluates MULTIPLE targeted therapies within a SINGLE disease histology,             |
|     stratified into distinct biomarker-defined sub-cohorts.                                       |
|   - Examples:                                                                                     |
|     * Lung-MAP (S1400): Squamous NSCLC stratified to MET, FGFR, PIK3CA, or PD-L1 targeted arms.   |
|     * ALCHEMIST: Adjuvant NSCLC screening for EGFR, ALK, and immunotherapy regimens.              |
|     * BATTLE: Biomarker-integrated trial in refractory NSCLC.                                     |
|                                                                                                   |
|   [3. PLATFORM TRIAL] (Multi-Arm Multi-Stage [MAMS] / Continuous Infrastructure)                  |
|   - Concept: Evaluates MULTIPLE therapies for a SINGLE disease perpetually. Investigational arms   |
|     enter and leave dynamically based on interim Bayesian adaptive rules against a shared control.|
|   - Examples:                                                                                     |
|     * I-SPY 2: Neoadjuvant therapy for high-risk early breast cancer.                             |
|     * STAMPEDE: Multi-arm randomized platform in advanced hormone-sensitive prostate cancer.      |
+---------------------------------------------------------------------------------------------------+
FeatureBasket TrialUmbrella TrialPlatform Trial
Tumor HistologiesMultiple tumor typesSingle tumor typeSingle tumor type (or defined disease stage)
Genomic AlterationsSingle genomic targetMultiple genomic targetsMultiple targets / phenotypic profiles
Investigational DrugsSingle drug (or combo)Multiple matched drugsMultiple drugs (dynamic entry/exit)
Control ArmTypically single-arm (no control)Common standard-of-care control armCommon shared control arm (MAMS)
Primary Regulatory GoalTissue-agnostic FDA approvalBiomarker-stratified approval within diseaseRapid screening and seamless Phase II/III transitions

4. Oncology Clinical Trial Endpoints & Response Criteria

Clinical trial endpoints are categorized based on whether they measure direct clinical benefit (how a patient feels, functions, or survives) or serve as surrogate intermediate markers.

Comparison of Oncology Trial Endpoints

EndpointDefinition & MeasurementAdvantagesLimitations & Biases
Overall Survival (OS)Time from randomization to death from any cause.Gold standard. Unambiguous, objective, directly reflects true clinical benefit; unaffected by measurement bias.Requires large sample size and prolonged follow-up; heavily confounded by post-progression effective crossover therapies.
Progression-Free Survival (PFS)Time from randomization to objective disease progression (RECIST) or death from any cause.Faster readout than OS; not confounded by subsequent line therapies; reflects tumor control.Susceptible to assessment bias and measurement interval bias; requires formal validation as surrogate for OS.
Time to Progression (TTP)Time from randomization to objective disease progression.Isolates tumor biology from non-cancer deaths.Deaths from non-cancer causes are censored, introducing competing risk bias. Rarely used today (PFS preferred).
Disease-Free Survival (DFS) / Event-Free Survival (EFS)Time from complete curative resection to disease recurrence or death from any cause (DFS/EFS).Standard primary endpoint in adjuvant and neoadjuvant curative trials.Requires long follow-up; definitions of "events" (second primary vs recurrence) must be strictly pre-specified.
Overall Response Rate (ORR)Proportion of patients achieving confirmed Complete Response (CR) + Partial Response (PR).Assessed in single-arm Phase II trials; rapid surrogate of drug antitumor activity.Does not measure disease stability; may not correlate with survival prolongation in slow-growing indolent tumors.
Pathologic Complete Response (pCR)Complete disappearance of invasive tumor cells in resected primary tissue and lymph nodes (ypT0/is ypN0).Early neoadjuvant surrogate for long-term EFS/OS in TNBC, HER2+ breast cancer, and resectable NSCLC.Requires surgical pathology resection; does not capture metastatic micro-dissemination directly.
Minimal Residual Disease (MRD)Detection of malignant cells below microscopic threshold ($<10^{-4}$ to $<10^{-6}$) via NGS or flow cytometry.Ultra-sensitive surrogate for PFS/OS in Multiple Myeloma, CLL, and B-ALL.Requires specialized bone marrow or peripheral blood assay standardization across clinical laboratories.

RECIST 1.1 vs iRECIST Criteria

Tumor response evaluation in solid tumor clinical trials relies on standardized radiographic criteria:

  • RECIST 1.1 (Response Evaluation Criteria in Solid Tumors):
    • Target Lesions: Up to a maximum of 5 total lesions (maximum 2 per organ), measured in the longest diameter (short axis for lymph nodes; normal $\text{LN} < 10\text{ mm}$, non-target $10\text{--}14\text{ mm}$, target $\ge 15\text{ mm}$).
    • Complete Response (CR): Disappearance of all target lesions; all pathological lymph nodes must reduce to $<10\text{ mm}$ short axis.
    • Partial Response (PR): At least a $\ge 30%$ decrease in the Sum of Diameters (SOD) of target lesions compared to baseline SOD.
    • Progressive Disease (PD): At least a $\ge 20%$ increase in the SOD of target lesions (with an absolute increase of $\ge 5\text{ mm}$) compared to the smallest SOD recorded (nadir), or the appearance of one or more new lesions.
    • Stable Disease (SD): Neither sufficient shrinkage to qualify for PR nor sufficient increase to qualify for PD.
  • iRECIST (Immune-Modified RECIST for Checkpoint Inhibitors):
    • Designed to account for pseudoprogression (transient immune cell infiltration into tumor causing radiographic enlargement or appearance of new lesions before subsequent tumor regression).
    • Introduces iUPD (immune Unconfirmed Progressive Disease): Upon initial documentation of progression, treatment may continue if the patient is clinically stable. Radiographic reassessment is mandated at 4–8 weeks.
    • iCPD (immune Confirmed Progressive Disease): Confirmed if the subsequent scan demonstrates further SOD increase ($\ge 5\text{ mm}$) or additional new lesions.

The Crossover Dilemma in Oncology Phase III Trials

When an experimental therapy demonstrates substantial efficacy, ethical considerations often permit patients in the control arm to "cross over" to receive the experimental drug upon disease progression. While crossover is ethically necessary, it severely confounds Overall Survival (OS) by artificially prolonging control arm survival, diluting the observed OS hazard ratio toward null ($HR \approx 1.0$), while PFS remains completely unconfounded.

Observed HROS1.00despite true antineoplastic efficacy!\text{Observed } HR_{\text{OS}} \to 1.00 \quad \text{despite true antineoplastic efficacy!}

Statistical adjustment methods (e.g., Rank-Preserving Structural Failure Time Models [RPSFTM], Two-Stage Estimation [TSE], or Inverse Probability of Censoring Weighting [IPCW]) are employed to estimate counterfactual OS without crossover.


5. Statistical Power, Sample Size Dynamics & Error Control

Hypothesis Testing & Error Types

+---------------------------------------------------------------------------------------------------+
|                         STATISTICAL HYPOTHESIS TESTING MATRIX                                     |
|                                                                                                   |
|                                       TRUE STATE OF NATURE                                        |
|   DECISION               Null Hypothesis (H0) True            Alternative Hypothesis (H1) True    |
|   ---------------------------------------------------------------------------------------------   |
|   Reject H0              TYPE I ERROR (alpha)                 CORRECT DECISION                    |
|   (Claim benefit)        False Positive Rate                  Statistical Power (1 - beta)        |
|                          (Standard: alpha = 0.05 two-sided)   (Standard: 80% to 90%)              |
|                                                                                                   |
|   Fail to Reject H0      CORRECT DECISION                     TYPE II ERROR (beta)                |
|   (No claim)             True Negative (1 - alpha)            False Negative Rate                 |
|                          (Confidence Level: 95%)              (Standard: beta = 0.10 to 0.20)     |
+---------------------------------------------------------------------------------------------------+

Event-Driven Power in Survival Trials (Schoenfeld's Formula)

In oncology trials with time-to-event endpoints (OS, PFS), statistical power does not depend on the total number of enrolled patients ($N$), but rather on the total number of observed events ($d$, deaths or progressions):

d=4×(Zα/2+Zβ)2(lnHR)2d = \frac{4 \times (Z_{\alpha/2} + Z_\beta)^2}{(\ln HR)^2}

Where:

  • $Z_{\alpha/2} = 1.96$ for two-sided $\alpha = 0.05$
  • $Z_\beta = 0.84$ for $80%$ power, or $1.28$ for $90%$ power
  • $HR$ = Target Hazard Ratio

Key Clinical Takeaway: If an oncology trial fails to observe the required number of target events (e.g., due to slower-than-expected disease progression or effective subsequent therapies), the trial remains underpowered regardless of how many hundreds of patients were enrolled.

Superiority vs Non-Inferiority (NI) Trial Designs

  • Superiority Design: Tests whether the experimental drug is clinically superior to the standard active control ($H_0: HR \ge 1.0$ vs $H_1: HR < 1.0$).
  • Non-Inferiority (NI) Design: Tests whether a novel agent (which offers other advantages: oral route, reduced toxicity, lower cost, shorter duration) is not unacceptably worse than the standard-of-care active comparator by more than a pre-specified margin $\Delta$.
    • Null Hypothesis: $H_0: HR \ge \Delta$ (Experimental drug is inferior by margin $\Delta$).
    • Alternative Hypothesis: $H_1: HR < \Delta$ (Non-inferiority demonstrated).
    • The Non-Inferiority Margin ($\Delta$): Must be clinically and statistically justified based on historical active control efficacy against placebo, ensuring that the new drug preserves a substantial fraction (e.g., $\ge 50%$) of the comparator effect.
    • Dataset Requirement: Non-inferiority trials mandate evaluating BOTH the Intention-to-Treat (ITT) and Per-Protocol (PP) populations. In NI trials, poor compliance, protocol deviations, and treatment dropouts bias the results toward equivalence (falsely favoring non-inferiority in ITT analyses), making the Per-Protocol analysis critically important.
Test Your Knowledge

A clinical investigator is designing a multicenter, single-arm Phase II trial evaluating a novel antibody-drug conjugate (ADC) in patients with HER2-low metastatic breast cancer. The investigator utilizes Simon's optimal two-stage design with a null response rate (p0) of 15% (unacceptable efficacy) and an alternative target response rate (p1) of 35% (promising efficacy), setting alpha = 0.05 and beta = 0.10 (90% power). In Stage 1, 19 evaluable patients are enrolled. The protocol specifies an early stopping boundary of r1 = 3 responses. At the completion of Stage 1, exactly 2 of the 19 patients achieve a confirmed partial response. What is the correct protocol-directed action?

A
B
C
D
Test Your Knowledge

A global oncology clinical trial protocol evaluates a novel highly selective RET inhibitor in patients whose tumors harbor RET gene fusions. Patients with RET fusion-positive non-small cell lung cancer, medullary thyroid cancer, papillary thyroid cancer, colorectal adenocarcinoma, and pancreatic ductal adenocarcinoma are enrolled across distinct histology cohorts receiving the same investigational RET inhibitor. How is this precision oncology master protocol classified?

A
B
C
D
Test Your Knowledge

A Phase III non-inferiority randomized trial compares an oral targeted doublet regimen against intravenous standard chemotherapy in metastatic colorectal cancer. The primary endpoint is overall survival (OS). The non-inferiority margin is pre-specified as a Hazard Ratio (HR) upper bound of delta = 1.25. The final intention-to-treat analysis yields a Hazard Ratio for OS of 1.08 with a 95% Confidence Interval of [0.92 to 1.21] (p for non-inferiority = 0.018; p for superiority = 0.42). What is the correct clinical and statistical interpretation of this trial result?

A
B
C
D
Test Your Knowledge

A randomized, open-label Phase III trial evaluates a novel third-generation tyrosine kinase inhibitor (TKI) versus standard chemotherapy in advanced non-small cell lung cancer. The primary endpoint is Progression-Free Survival (PFS), and the key secondary endpoint is Overall Survival (OS). The study reports a dramatic and statistically significant PFS improvement in the experimental arm (Median PFS: 18.9 months vs 5.4 months; HR 0.42; 95% CI: 0.34–0.52; p < 0.0001). However, the final Overall Survival analysis demonstrates no statistically significant difference between arms (Median OS: 31.8 months vs 30.2 months; HR 0.94; 95% CI: 0.78–1.14; p = 0.54). Further review reveals that 82% of patients randomized to the control chemotherapy arm crossed over to receive the experimental TKI upon disease progression. How should the clinical oncology pharmacist interpret these findings?

A
B
C
D