5.4 Step 3: The Critical Interpretation & Appraisal Process

Key Takeaways

  • Step 3 of CHD's eight-step EBD process—critically interpret relevant evidence—filters search results into credible, relevant evidence before concepts and hypotheses are developed.
  • EDAC study materials evaluate existing evidence on three criteria: relevance, validity, and reliability.
  • A seven-question appraisal checks the research question, study design, sample size and power, measurement, confounders, stated limitations, and causal claims.
  • Establishing causation requires temporal precedence, covariation of cause and effect, and elimination of plausible alternative explanations.
  • A statistically significant result (for example, p < 0.05) can still be too small to matter, so effect size and practical significance must be judged separately.
Last updated: September 2026

Step 3: The Critical Interpretation & Appraisal Process

The Methodological Bridge: In CHD's 8-step Evidence-Based Design process, Step 3 (Critically Interpret Evidence) serves as the crucial intellectual filter between information gathering and design execution. Step 2 searches for and gathers the raw evidence; Step 3 subjects that evidence to rigorous critical appraisal; only then can Step 4 create and innovate design concepts, followed by Step 5 developing testable hypotheses.

Without Step 3, healthcare design teams risk committing the fatal error of premature design translation—taking an unvalidated white paper, an underpowered correlation, or an irrelevant pediatric study and embedding it into the concrete foundations of an adult surgical hospital. Critical interpretation is the discipline that ensures only credible, valid, and generalizable evidence shapes the built environment.


The EDAC 8-Step Process Context

To master Step 3, candidates must understand its precise position within the complete EDAC project lifecycle:

Step 1: Define evidence-based goals and objectives
  │
Step 2: Find sources for relevant evidence
  │
Step 3: CRITICALLY INTERPRET RELEVANT EVIDENCE ◄── [The Scientific Filter]
  │
Step 4: Create and innovate EBD concepts
  │
Step 5: Develop a hypothesis
  │
Step 6: Collect baseline performance measures
  │
Step 7: Monitor implementation of design and construction
  │
Step 8: Measure post-occupancy performance results

In Step 1, the interdisciplinary team establishes baseline organizational aspirations (e.g., "Reduce hospital-acquired central-line bloodstream infections by 30%"). In Step 2, researchers conduct exhaustive literature searches across academic databases. In Step 3, the team critically dissects every retrieved study to determine:

  • Is the research scientifically sound?
  • Are the findings relevant to our specific clinical population and facility scale?
  • Can we translate these empirical conclusions into concrete, measurable architectural specifications?

The 7-Question Critical Appraisal Methodology

EDAC study materials name three criteria for evaluating existing evidence: relevance, validity, and reliability. The seven questions below are a practical way to apply those criteria to any research publication, post-occupancy evaluation, or vendor report:

1. What was the exact research question or guiding hypothesis?

  • Is the research question clearly stated, focused, and testable?
  • Does the hypothesis articulate a specific directional relationship between an environmental feature (independent variable) and a clinical, operational, or behavioral outcome (dependent variable)?
  • Red Flag: Vague, exploratory papers that lack an explicit research question and wander into post-hoc data dredging.

2. Where does the study design fall on the hierarchy of evidence?

  • Is the study a Level I systematic review, a Level II RCT, a Level III quasi-experimental study, a Level IV correlational study, a Level V descriptive POE, or Level VI expert opinion?
  • Was there an active control or comparison group, or is the study an uncontrolled before-and-after observation?
  • Red Flag: Uncontrolled single-facility case studies claiming definitive clinical proof.

3. Was the sample size adequate and statistically powered?

  • Did the study include an adequate sample size (N) to detect a statistically significant difference without committing a Type II error (falsely concluding an environmental feature has no effect when it actually does)?
  • Did the authors report a statistical power analysis (commonly targeting power, 1 − β, of at least 0.80 at α = 0.05)?
  • Red Flag: Studies evaluating patient falls across only 15 patient-days or surveying only 8 staff members, where statistical noise overwhelms true effects.

4. How were independent and dependent variables operationalized?

  • Were environmental variables measured using calibrated, objective physical instruments (e.g., continuous Class 1 sound meters recording LAeq, digital lux meters, particle counters, automated airflow hood sensors) or subjective, ambiguous descriptors (e.g., "the room felt brightly lit and calm")?
  • Were clinical outcomes measured using standardized, validated definitions (e.g., CDC National Healthcare Safety Network [NHSN] criteria for infections; National Database of Nursing Quality Indicators [NDNQI] fall definitions; validated HCAHPS surveys)?
  • Red Flag: Studies where environmental conditions are unquantified assumptions rather than measured data.

5. What confounding variables existed, and how were they controlled?

  • Did the researchers identify concurrent operational changes (EHR rollouts, staffing shifts, protocol updates)?
  • Did the study employ concurrent matched comparison units, interrupted time-series longitudinal tracking, or multivariable statistical adjustments (e.g., ANCOVA, logistic regression controlling for patient acuity)?
  • Red Flag: Studies asserting that architectural renovations caused clinical improvements without tracking or controlling for clinical protocol changes.

6. What explicit limitations and biases did the authors acknowledge?

  • Do the authors present a transparent, critical "Limitations" section discussing sample constraints, missing data, potential confounding factors, and generalizability boundaries?
  • Red Flag: Any study that claims 100% conclusive results without acknowledging a single methodological limitation or risk of bias. A missing limitations section is a warning sign of weak scholarship or marketing.

7. Did the study establish true causality or merely statistical correlation?

  • Did the study design satisfy the rigorous scientific prerequisites required to prove causality, or did it merely demonstrate that two variables occurred at the same time?
  • Red Flag: Authors who conclude that "installing artwork causes faster healing" based purely on a cross-sectional correlational survey.

The Core Imperative: Distinguishing Correlation from Causation

A central skill in Step 3 is differentiating statistical correlation from causation.

The Three Prerequisites for Establishing Causality

To legitimately claim that an architectural design feature caused a clinical or operational outcome, three conditions must be met simultaneously:

  1. Temporal Precedence: The cause must precede the effect in time. The physical environmental modification must be fully operational before the change in clinical outcome occurs.
  2. Covariation of Cause and Effect: When the environmental feature is present (or altered), the clinical outcome changes systematically; when the environmental feature is absent (or reversed), the outcome does not change (or reverts).
  3. Non-Spuriousness (Elimination of Plausible Alternative Explanations): The researcher must prove that no third, unmeasured confounding variable (e.g., new leadership, technological upgrades, staffing adjustments) produced the observed change.
   CORRELATION DOES NOT EQUAL CAUSATION
   
   Correlation: Variable A and Variable B move together.
   Example: Facilities with decentralized nurse stations also have higher patient satisfaction.
   
   Spurious Third Variable:
   Perhaps facilities with decentralized stations ALSO invest in 30% higher nurse staffing.
   The staffing difference, not the desk layout, might explain the higher satisfaction.

Statistical Significance vs. Clinical & Operational Significance

When critically interpreting empirical data, EDAC professionals must distinguish between two fundamentally different concepts:

Statistical Significance (p-value)

  • Definition: The p-value is the probability of observing a difference at least as large as the one found if the null hypothesis (no real difference) were true.
  • Standard Threshold: Commonly p < 0.05 (EDAC study materials describe rejecting the null hypothesis when p is less than 5%), sometimes p < 0.01.
  • The Sample Size Trap: In massive datasets (e.g., 100,000 patient-days or 50,000 acoustic sensor readings), virtually any microscopic, trivial variation becomes "statistically significant" (p < 0.001). A large sample size detects noise with extreme statistical certainty.

Clinical and Operational Significance (Effect Size)

  • Definition: Reflects the magnitude, practical real-world relevance, and clinical impact of the outcome change.
  • Metrics: Quantified using standardized effect sizes such as Cohen's d (standardized mean difference), Odds Ratios (OR), Relative Risk Reduction (RRR), or return on investment (ROI).
  • The Reality Check:
    • A study demonstrates that an expensive specialized acoustic plaster reduces average ambient ICU noise by 0.8 decibels (LAeq), achieving high statistical significance (p = 0.004) across 20,000 hours of recording.
    • However, a change of about 1 dB is barely detectable even under laboratory conditions (about 3 dB is generally the smallest change people notice in everyday settings), and ICU delirium and sleep disruption rates remained unchanged.
    • Conclusion: The plaster achieved statistical significance but showed little practical significance. Unless other benefits justify it, a large cost premium for this material would be hard to defend.

Reconciling Contradictory Research Findings

In healthcare architecture, published literature often presents contradictory conclusions. For example, Study A may report that decentralized nurse stations significantly reduce nurse travel fatigue, while Study B reports that decentralized stations increase nurse burnout and feelings of professional isolation.

When confronting conflicting literature, an EBD team does not flip a coin or pick the study they like best. Instead, they execute a methodological reconciliation audit:

  1. Compare Clinical Acuity and Specialty: Did Study A examine a 32-bed post-surgical elective recovery unit where stable patients require routine checks, while Study B examined an acute pediatric intensive care unit where complex titrations require constant two-nurse verification?
  2. Examine Architectural Floor-Plate Geometry: Did Study A provide visual sightlines between decentralized alcoves and a central resource hub, while Study B placed alcoves in deep, enclosed architectural recesses that completely blocked peer visibility?
  3. Audit Technology Integration: Did Study A equip nurses with hands-free VoIP communication badges and real-time locator badges, while Study B forced nurses to walk to a central station to find an available phone?
  4. Synthesize Hybrid Solutions: By critically understanding why the findings diverged, the design team can synthesize an optimized, hybrid design concept—such as decentralized charting alcoves paired with a centralized team communication core—satisfying both operational needs.

Translating Appraised Evidence: The Design Criteria Matrix

The culmination of Step 3 is synthesizing all critically interpreted research into a structured, operational document: the Design Criteria Matrix (or Matrix of Evidence). This matrix bridges the gap between scientific literature and architectural schematic design, providing the foundational inputs for Step 4 (Concept Innovation) and Step 5 (Hypothesis Formulation).

Example Design Criteria Matrix for an Acute Care Inpatient Unit (Illustrative)

The rows show how appraised evidence is recorded; strength ratings and targets must come from your own appraisal and baseline data.

Clinical / Operational GoalEvidence Summary & Appraised StrengthArchitectural Parameter / Design SolutionProposed Outcome Metric (Steps 5, 6, 8)Confidence / Risk Level
Reduce hospital-acquired infectionsResearch reviews associate single-patient rooms and accessible hand-hygiene facilities with lower transmission risk; most studies are observational or before-and-afterSingle-patient rooms; hand-hygiene sink visible on entry; cleanable surfacesTargeted HAI rates per 1,000 patient daysModerate–High
Reduce inpatient falls with injuryStudies of bed-to-toilet paths and patient visibility are mixed and context-dependentClear, supported path from bed to toilet; visibility from staff work areasFalls and falls with injury per 1,000 patient days (standard definitions)Moderate
Reduce sleep disruptionStudies link noise peaks to awakenings; sound absorption lowers reverberation and noise levelsSound-absorbing ceilings; alarm management; door sealsSound levels (Leq, Lmax); patient-reported sleep; sedative useModerate
Improve nurse efficiencyTime-and-motion studies (e.g., Hendrich et al., 2008) show much nursing time goes to non-bedside work and walkingDecentralized supplies and medications near patient rooms; hybrid work areasWalking distance; time in patient roomsModerate

Final Checklist: Red Flags in Research Appraisal for EDAC Candidates

Before allowing any empirical study to influence schematic design or project hypotheses, EDAC candidates must verify that the study passes this final appraisal checklist:

  • No Commercial Conflict of Interest: Was the research funded or authored by a manufacturer whose product was tested?
  • Validated Measurement Tools: Were physical and clinical metrics gathered using standardized, calibrated, objective instruments rather than unverified questionnaires?
  • Adequate Observation Window: Did post-occupancy tracking continue for at least 6 to 12 months after move-in to allow operational workflows to stabilize and novelty effects to fade?
  • Transparent Reporting of Limitations: Did the authors explicitly detail study constraints, unmeasured confounders, and generalizability boundaries?
  • Distinction Between Correlation and Causation: Did the authors refrain from claiming causal architectural proof when utilizing non-experimental correlational designs?

[!CAUTION]

EXAM TRAP: Skipping Step 3 in the EBD Lifecycle

A typical scenario describes a project team that establishes evidence-based goals in Step 1, conducts a thorough literature search in Step 2, and immediately begins sketching architectural concepts and building full-scale mockups in Step 4.

What critical step was skipped? The team skipped Step 3 (Critically Interpret Evidence). Without critically appraising the gathered studies for internal validity, sample size, confounders, and generalizability, the team risks embedding flawed, unreplicable, or irrelevant research into physical mockups and permanent construction. Skipping Step 3 invalidates the integrity of the entire EBD methodology.

Loading diagram...
Step 3: The Critical Interpretation & Appraisal Workflow
Test Your Knowledge

In The Center for Health Design's 8-step Evidence-Based Design process, where does the critical interpretation and appraisal of evidence occur, and what is its immediate purpose?

A
B
C
D
Test Your Knowledge

An architectural team reviews a study reporting a strong statistical correlation (r = 0.78, p < 0.01) between the installation of decentralized nurse workstations and an increase in patient satisfaction scores. The lead designer asserts this proves decentralized workstations directly caused higher patient satisfaction. What methodological principle must the EDAC consultant highlight to challenge this conclusion?

A
B
C
D
Test Your Knowledge

A research paper indicates that a specialized acoustic ceiling tile reduced ambient peak noise levels in an intensive care unit by 1.1 decibels, achieving statistical significance at p = 0.03 due to a massive sample size of 50,000 acoustic measurements. However, patient sleep disruption scores and sedative medication usage showed zero measurable difference. When translating these findings during Step 3 into design recommendations, how should the team interpret this evidence?

A
B
C
D