5.1 The Hierarchy of Evidence: Systematic Reviews to Expert Opinion

Key Takeaways

  • An evidence hierarchy ranks research designs by susceptibility to bias; this guide uses a working six-level version, and published hierarchies differ in labels and levels.
  • Systematic reviews and meta-analyses follow reproducible search and appraisal protocols, which distinguishes them from narrative literature reviews.
  • Randomized controlled trials are rarely feasible for whole buildings but are more practical for smaller, modular interventions such as lighting or sound conditions.
  • Quasi-experimental studies with comparison units or interrupted time series are among the strongest designs usually feasible for healthcare facilities.
  • EDAC study materials note that a quasi-experimental study with a control group is more credible than a correlational study when evaluating design innovations.
Last updated: September 2026

The Hierarchy of Evidence: Systematic Reviews to Expert Opinion

Core Concept: In Evidence-Based Design (EBD), not all research carries equal weight. The Hierarchy of Evidence is a structured heuristic adapted from Evidence-Based Medicine (EBM) that ranks research designs according to their susceptibility to bias and their internal validity—the degree of certainty that an environmental design intervention directly produced the observed clinical, operational, or safety outcome.

Healthcare design decisions routinely commit tens or hundreds of millions of dollars to physical environments that will operate for 30 to 50 years. Making these irreversible capital investments based on flawed data, vendor marketing claims, or subjective architectural anecdotes exposes healthcare systems to severe financial, operational, and clinical risks. For this reason, EDAC-certified professionals must critically evaluate every empirical claim against its relative position on the evidence hierarchy.


The Hierarchy of Evidence Adapted for the Built Environment

In clinical medicine, the Oxford Centre for Evidence-Based Medicine (CEBM) hierarchy places double-blind randomized placebo-controlled trials (RCTs) near the apex. However, physical buildings cannot be blinded, nor can patients be randomized to receive "placebo" structural walls or substandard emergency departments. Consequently, healthcare design researchers adapt the hierarchy to the physical and ethical realities of built-environment research. The six-level version below is a working framework used in this guide; published hierarchies differ in how many levels they use and how they label them.

  • Level I: Systematic Reviews & Meta-Analyses (highest inferential strength; lowest bias risk)
  • Level II: Randomized Controlled Trials (RCTs) (rare in macro-architecture; common in micro-interventions)
  • Level III: Quasi-Experimental Studies (practical gold standard: concurrent control units & ITS)
  • Level IV: Non-Experimental Correlational & Cohort Studies (case-control, cross-sectional, and space syntax studies)
  • Level V: Descriptive Studies & Uncontrolled POEs (benchmarking surveys, qualitative focus groups)
  • Level VI: Expert Opinion, Case Studies & Vendor Claims (anecdotal reports, architectural awards, white papers)

Detailed Analysis of the Six Evidence Levels

Level I: Systematic Reviews and Meta-Analyses

A systematic review is a comprehensive, rigorous synthesis of all available empirical studies addressing a specific, focused research question. Unlike traditional narrative literature reviews, a systematic review follows a predefined, reproducible protocol (such as PRISMA—Preferred Reporting Items for Systematic Reviews and Meta-Analyses) that includes:

  • Explicit inclusion and exclusion criteria.
  • Exhaustive searches across multiple multidisciplinary databases (e.g., PubMed, CINAHL, PsycINFO, Avery Index).
  • Dual-independent reviewer screening to minimize selection bias.
  • Standardized critical appraisal scoring to weight studies by methodological quality.

When individual studies report quantitative metrics with sufficient homogeneity, researchers conduct a meta-analysis—a statistical pooling of effect sizes (e.g., Cohen's d, odds ratios, or relative risk reductions) that increases statistical power and produces a more precise pooled estimate of the effect than any single study.

Landmark Level I Syntheses in EBD

  • Ulrich, Zimring, and colleagues (2004 report for CHD; 2008 review in HERD): Extensive reviews of empirical studies linking hospital design features—single-patient rooms, noise, light, layout, and more—to patient and staff outcomes. They are widely cited literature reviews, but they are not quantitative meta-analyses, so read them as structured summaries of the evidence rather than pooled effect estimates.
  • CHD research summaries and issue briefs: Topic-focused summaries of healthcare design research.
  • Cochrane and other formal systematic reviews: Reviews on related clinical and environmental questions that follow explicit protocols.

Inferential Rigor: Highest when well conducted. Systematic reviews reduce the influence of any single facility's quirks, though their conclusions are only as good as the studies they include.


Level II: Randomized Controlled Trials (RCTs)

A Randomized Controlled Trial (RCT) is a true experimental design characterized by three indispensable pillars:

  1. Random Allocation: Participants (or clinical units) are assigned to the intervention or control group purely by chance, ensuring that confounding variables (e.g., patient age, comorbidity, socioeconomic status) are evenly distributed across groups.
  2. Active Manipulation: The researcher directly manipulates the independent variable (the environmental feature) while holding all other factors constant.
  3. Control Group: An identical cohort receiving standard treatment (or placebo) for concurrent comparison.

The Architectural Constraint and Micro-Interventions

Executing an RCT for macro-architectural layout (such as building two identical $250-million hospitals and randomizing trauma patients to a radial pod vs. a double-loaded corridor) is physically impossible, financially absurd, and ethically unacceptable under bioethical codes (such as the Belmont Report).

However, Level II RCTs are successfully executed in micro-environmental interventions:

  • Lighting Trials: Randomizing preterm infants in a neonatal intensive care unit (NICU) to cycled day-night lighting versus near-constant lighting, then measuring outcomes such as weight gain and length of stay (cycled lighting has been studied in randomized trials).
  • Acoustic Sound-Masking Trials: Randomizing adjacent medical-surgical rooms to active speech-privacy sound-masking spectra vs. inactive baseline speakers, evaluating polysomnographic sleep continuity and speech intelligibility index (SII).
  • Virtual Reality (VR) Simulation Trials: Randomizing clinical staff or surgical candidates to immersive natural biophilic restorative environments versus abstract geometrical environments, measuring salivary cortisol, skin conductance, and heart rate variability (HRV).

Inferential Rigor: Very high for the specific intervention tested, when randomization and follow-up are sound.


Level III: Quasi-Experimental Studies (The Practical Gold Standard in EBD)

A quasi-experimental study incorporates experimental manipulation and structured comparison, but lacks true random assignment of subjects to physical settings. Because hospital administrators cannot randomly assign which acutely ill patient receives a renovated room, quasi-experiments represent the practical gold standard for evaluating healthcare facilities.

Primary Quasi-Experimental Designs in Healthcare Architecture

  1. Non-Equivalent Control Group Pretest-Posttest Design:

    • A hospital renovates Unit 4A (installing decentralized nurse stations, high-NRC acoustic ceiling tiles, and solid-surface seamless millwork).
    • Unit 4B (an identical, unrenovated unit on the floor below with the same clinical specialty and patient acuity) serves as the concurrent control.
    • Baseline metrics (fall rates, nurse travel distance, hospital-acquired infections) are collected in both units simultaneously before renovation, and tracked in both units for 12 months post-occupancy.
    • Because Unit 4B experiences the same seasonal illness waves, hospital administrative leadership, and EHR updates, it controls for secular trends.
  2. Interrupted Time-Series (ITS) Design:

    • Researchers collect repeated outcome measures at regular, equal intervals (e.g., monthly fall rates) over an extended timeframe: 12 to 24 months pre-intervention and 12 to 24 months post-intervention.
    • Statistical modeling analyzes both the step change (immediate drop or spike upon occupancy) and the trend slope change (sustained improvement over time), separating true environmental impact from natural baseline trajectories.
  3. Stepped-Wedge Cluster Design:

    • A multi-phase architectural rollout where different hospital wings receive design upgrades sequentially over time, with each wing acting as its own pre-renovation control before transitioning to the intervention phase. (When the order of rollout is randomly assigned, the design becomes a stepped-wedge cluster randomized trial.)

Inferential Rigor: High when well designed. Controls for many threats to internal validity when comparison groups are well matched, although unmeasured differences can remain.


Level IV: Non-Experimental, Correlational, Cohort & Case-Control Studies

Level IV studies observe and quantify existing relationships between environmental variables and outcomes without direct experimental manipulation by the researcher.

  • Prospective & Retrospective Cohort Studies: Tracking defined patient groups exposed to different existing environments over time. For example, Roger Ulrich's classic 1984 Science study was a retrospective matched-pair study: 46 cholecystectomy patients paired on factors such as sex, age, smoking, obesity, prior hospitalization, and floor level, comparing patients whose windows faced trees with patients whose windows faced a brick wall.
  • Case-Control Studies: Identifying patients who experienced an adverse event (e.g., 100 inpatient falls in the bathroom) and comparing them to matched controls who did not fall, isolating physical environmental differences such as door swing radius, grab bar orientation, or floor slip-resistance ratings.
  • Cross-Sectional Correlational Surveys: Measuring environmental features and occupant responses simultaneously across multiple sites (e.g., calculating travel distances across 40 hospital floor plans using space syntax analysis and correlating them with nurse burnout scores).

Inferential Rigor: Moderate. Can demonstrate associations, but cannot rule out unmeasured confounding factors. EDAC study materials note, for example, that a quasi-experimental study with a control group is more credible than a correlational study for evaluating a design innovation's effect on noise.


Level V: Descriptive Studies & Uncontrolled Post-Occupancy Evaluations (POEs)

Level V studies document what exists in a single facility or cohort without using a comparison or control group:

  • Single-Site Post-Occupancy Evaluations (POEs): Standard post-move surveys measuring occupant satisfaction, perceived thermal comfort, or acoustic privacy 6 to 12 months after moving into a new building, without pre-move baseline metrics or an unrenovated comparison wing.
  • Benchmarking Surveys: Cross-sectional questionnaires sent to hundreds of facility managers documenting square footage allocations or finish specifications.
  • Qualitative Ethnographic Observations: Structured behavioral mapping, shadow-tracking of clinicians, or patient focus groups documenting lived experiences and operational workarounds.

Inferential Rigor: Low to Moderate. Extremely valuable for generating hypotheses, uncovering unexpected operational bottlenecks, and refining functional programming, but structurally incapable of establishing causal attribution.


Level VI: Expert Opinion, Case Studies & Vendor Claims

The base of the pyramid represents information derived from subjective authority, commercial interests, or unvalidated individual experiences:

  • Consensus Committee Guidelines & Expert Opinion: White papers issued by professional architectural associations, advisory panels, or thought leaders based on collective clinical or architectural experience rather than formal empirical testing.
  • Architectural Monographs & Design Award Submissions: Descriptive publications showcasing aesthetic elegance, spatial volume, and form, often written before the building is occupied or without verifying clinical outcome metrics.
  • Vendor-Sponsored White Papers & Commercial Claims: Product literature published by building product manufacturers claiming proprietary performance benefits (e.g., antimicrobial copper alloys, specialized acoustic baffles, or ergonomic patient recliners) tested only under ideal laboratory conditions without peer-reviewed field validation.

Inferential Rigor: Lowest. Highly prone to confirmation bias, commercial conflict of interest, and survivorship bias.


Comprehensive Comparative Matrix of Evidence Levels

LevelStudy Design TypologyTypical EBD Healthcare ApplicationInferential RigorSusceptibility to BiasCapital Justification Strength
Level ISystematic Reviews & Meta-AnalysesProtocol-driven reviews on a focused design questionVery HighVery LowStrongest: Preferred foundation for major decisions when available
Level IIRandomized Controlled Trials (RCTs)Cycled NICU lighting trials; randomized room-condition experimentsHigh to Very HighLowStrong: Best for testing specific, modular interventions
Level IIIQuasi-Experimental (Controlled Pre/Post, ITS)Renovated vs. unrenovated comparison wings; 24-month interrupted time-seriesHighLow to ModerateStrong: Practical gold standard for hospital architecture
Level IVNon-Experimental (Cohort, Case-Control, Correlational)Matched cohort room orientation (Ulrich 1984); space syntax travel distanceModerateModerateSubstantial: Solid empirical rationale; requires context matching
Level VDescriptive Studies, Uncontrolled POEsSingle-site post-move satisfaction surveys; nurse focus groups; behavioral mappingLow to ModerateHighSupplementary: Valuable for hypothesis generation, not proof
Level VIExpert Opinion, Design Awards, Vendor ClaimsManufacturer white papers on finishes; architect design portfolios; consensus reportsLowestVery HighInsufficient: Cannot justify capital investments without peer-reviewed data

Critical Rules for EDAC Candidates Evaluating Empirical Claims

When evaluating published research during the predesign and programming phases, EDAC candidates must apply three foundational rules:

  1. The Weight of Evidence Principle: A single Level IV study showing a positive result does not override an established Level I systematic review showing no effect or contradictory outcomes. Always weight decisions toward the apex of the hierarchy.
  2. The Independence Standard: Vendor-sponsored white papers (Level VI) should never be accepted at face value. A material specification (such as a costly self-disinfecting surface coating) must be substantiated by independent, third-party, peer-reviewed clinical field trials (Level II, III, or IV).
  3. Methodological Congruence: The evidence sought should match the scale and financial risk of the decision. Minor aesthetic finish selections may reasonably proceed on descriptive evidence and consensus, but irreversible, high-cost decisions (e.g., unit layout and nursing station configuration) warrant the strongest evidence available—ideally Level I–III—supplemented by mock-ups and local data when higher-level studies do not exist.

[!CAUTION]

EXAM TRAP: Confusing Design Awards and Vendor White Papers with Empirical Evidence

A common scenario describes a situation where a hospital committee selects a complex, expensive layout because it won an American Institute of Architects (AIA) design award, or selects an expensive interior finish based on a manufacturer's glossy white paper claiming a 50% infection drop.

Neither constitutes valid empirical evidence. Design awards celebrate aesthetic composition, formal massing, and architectural intent—rarely clinical outcomes. Manufacturer white papers are Level VI marketing materials loaded with commercial bias. Credible justification for major healthcare investments should rest on peer-reviewed empirical research, with lower-level sources used to generate ideas and hypotheses.

Loading diagram...
The Hierarchy of Evidence in Evidence-Based Design
Test Your Knowledge

A hospital design team planning a new neonatal intensive care unit (NICU) wants to implement a circadian lighting system. They locate a study where 60 premature infants were randomly allocated to either a standard static lighting nursery or a dynamic circadian lighting nursery with automated spectral tuning, measuring weight gain and length of stay. What level of evidence does this represent, and why?

A
B
C
D
Test Your Knowledge

A flooring manufacturer provides an architectural team with an internal white paper claiming that their patented rubber flooring product reduces nurse joint fatigue by 45% and eliminates 99.9% of surface pathogens within two hours. Based on the EBD hierarchy of evidence, how should an EDAC-certified professional evaluate and apply this claim?

A
B
C
D
Test Your Knowledge

What fundamental methodological characteristic distinguishes a Level I systematic review from a traditional narrative literature review or architectural monograph?

A
B
C
D