17.1 Evidence-Based Practice: Clinical Questions & Critical Appraisal
Key Takeaways
- Domain 4 Task 1 requires methods for locating, reviewing, interpreting, and critically appraising scholarly research to guide practice-relevant decision-making, naming defining a clinical question and determining the clinical bottom line as examples.
- Evidence-based practice integrates three sources: the best available research evidence, the practitioner's clinical expertise, and the client's values, preferences, and circumstances — research alone is not evidence-based practice.
- PICO structures a searchable question: Population, Intervention, Comparison, and Outcome.
- In the five-level hierarchy used in occupational therapy literature, Level 1 comprises systematic reviews, meta-analyses, and randomized controlled trials, descending to Level 5 case reports and expert opinion.
- Statistical significance is not clinical significance: a result with a p-value below 0.05 can still fall below the minimal clinically important difference and mean nothing to the client.
Evidence-Based Practice: Clinical Questions & Critical Appraisal
Domain 4, Task 1 opens with methods for locating, reviewing, interpreting, and critically appraising scholarly research to guide practice-relevant decision-making, naming defining a clinical question and determining the clinical bottom line as its examples. This content is easy to skip and reliably tested.
1. What Evidence-Based Practice Actually Is
Evidence-based practice integrates three sources, and the exam tests all three:
- Best available research evidence
- Clinical expertise — the practitioner's accumulated skill and reasoning
- Client values, preferences, and circumstances
A therapist who applies a Level 1 protocol to a client who will not do it, cannot afford it, or does not value the outcome is not practicing evidence-based care. Conversely, "this is how we've always done it" is expertise without evidence. The keyed answer almost always integrates all three.
2. The Five Steps
| Step | What You Do |
|---|---|
| 1. ASK | Convert a clinical uncertainty into a focused, answerable question using PICO |
| 2. ACQUIRE | Search efficiently in the right databases |
| 3. APPRAISE | Critically evaluate the quality, size, and applicability of what you found |
| 4. APPLY | Integrate the evidence with your expertise and the client's values and circumstances |
| 5. ASSESS | Evaluate the outcome for this client and for your own practice |
PICO
| Element | Question | Example |
|---|---|---|
| P — Population / Problem | Who is the client? | Adults within 6 months of stroke with moderate upper extremity hemiparesis |
| I — Intervention | What are you considering? | Constraint-induced movement therapy |
| C — Comparison | Compared to what? | Conventional task-oriented training |
| O — Outcome | What result matters? | Improved performance in self-feeding and dressing |
Assembled: "In adults within six months of stroke with moderate upper extremity hemiparesis, does constraint-induced movement therapy, compared with conventional task-oriented training, improve independence in self-feeding and dressing?" That question is searchable. "Is CIMT good?" is not.
Where to Search
Peer-reviewed databases and pre-appraised sources: the American Journal of Occupational Therapy, PubMed and MEDLINE, CINAHL, the Cochrane Library, OTseeker, PEDro, and AOTA's evidence-based practice resources including its practice guidelines and critically appraised topics. Pre-appraised sources save enormous time — a systematic review has already done the appraisal work.
3. Levels of Evidence
Occupational therapy literature commonly uses a five-level hierarchy:
| Level | Design | Strength |
|---|---|---|
| Level 1 | Systematic reviews, meta-analyses, and randomized controlled trials | Strongest control of bias |
| Level 2 | Non-randomized studies with two groups — cohort and case-control designs | Group comparison without randomization |
| Level 3 | Non-randomized studies with one group — pre-test/post-test designs | No comparison group; change may reflect natural recovery |
| Level 4 | Descriptive studies including single-subject designs and case series | Detailed but limited generalizability |
| Level 5 | Case reports and expert opinion | Weakest; hypothesis-generating |
Critical nuance: level is not the same as quality. A poorly conducted randomized trial with 12 participants and 40 percent attrition is weaker evidence than a rigorous, well-powered cohort study. And a single-subject design — despite sitting at Level 4 — is often the most appropriate design for a low-incidence population and produces genuinely useful practice information.
Designs Worth Recognizing
- Randomized controlled trial: participants randomly assigned; the design best able to support a causal claim.
- Cohort: groups followed forward in time; may be prospective or retrospective.
- Case-control: starts from the outcome and looks backward for exposures; efficient for rare outcomes.
- Cross-sectional: a snapshot at one time point; describes prevalence, not causation.
- Single-subject designs (ABA, ABAB, multiple baseline): repeated measurement of one participant across withdrawn and reintroduced conditions.
- Qualitative designs (phenomenology, grounded theory, ethnography): answer how and why questions about lived experience — the natural fit for occupational meaning, identity, and client experience. Qualitative work is not "weak quantitative work"; it answers a different question.
- Mixed methods: combines both to answer both kinds of question.
4. Statistics the Practitioner Must Interpret
| Concept | Meaning | Practical Reading |
|---|---|---|
| p-value | Probability of the observed result if there were truly no effect | p < 0.05 is the conventional significance threshold. It says nothing about the size or importance of the effect |
| Confidence interval | The range of plausible values for the true effect | A 95% interval crossing zero (or 1 for a ratio) indicates a non-significant result. Narrow intervals mean precision |
| Effect size (e.g., Cohen's d) | Magnitude of the difference | Roughly 0.2 small, 0.5 medium, 0.8 large |
| Minimal detectable change | Smallest change exceeding measurement error | Below this, change is noise |
| Minimal clinically important difference | Smallest change the client perceives as beneficial | This is the number that matters to the client |
| Number needed to treat | How many clients must receive the intervention for one additional good outcome | Lower is better |
| Power | Ability to detect an effect that exists | Underpowered studies produce false negatives |
| Intention-to-treat analysis | Participants analyzed in their assigned group regardless of adherence | Preserves the benefit of randomization; more conservative and more realistic |
The distinction the exam builds items around: statistically significant and clinically meaningful are different claims. A study of 4,000 people can produce p < 0.001 for a 2-point change that no client would notice.
5. Appraising Quality
Internal validity asks whether the study's design supports its causal claim. External validity asks whether the findings apply to your client.
Common threats to watch for:
- Selection bias — groups differed at baseline.
- Attrition — differential dropout, especially over 20 percent.
- Maturation — natural recovery mistaken for treatment effect; a serious problem in acute stroke and pediatric studies without a control group.
- History — an outside event during the study period.
- Testing effects — improvement from repeated exposure to the measure.
- Hawthorne effect — participants change because they are being observed.
- Lack of blinding — unblinded assessors inflate effects, especially with subjective outcomes.
- Small or homogeneous samples — limits generalization.
Applicability questions: Were participants like my client in age, diagnosis, severity, and setting? Was the dose deliverable in my setting? Were the outcomes ones my client cares about? Do the benefits outweigh the harms and cost for this person?
Structured appraisal tools include the Critical Appraisal Skills Programme checklists and the PEDro scale.
6. The Clinical Bottom Line
The blueprint names this explicitly. A clinical bottom line is the short, practice-facing translation of the evidence — what you would tell a colleague in the hallway. A usable one states:
- What the evidence supports — the intervention, the population, and the outcome.
- How strong it is — the level, quantity, and consistency of the evidence.
- What it does not say — the boundaries and unanswered questions.
- What you will do for this client, given their values and circumstances.
Example: "Multiple Level 1 studies support constraint-induced movement therapy for improving upper extremity function in clients with at least 10 degrees of active wrist and finger extension. My client meets the motor criteria but works full time and cannot commit to six hours a day of restraint. I will implement a modified protocol with two to three hours of daily restraint and structured practice, and reassess with the Canadian Occupational Performance Measure at four weeks."
That paragraph integrates evidence, expertise, and client circumstances — which is precisely what evidence-based practice means.
7. Making It Sustainable
The universal barriers are time, database access, and appraisal confidence. The practical responses:
- Start with pre-appraised evidence — practice guidelines, systematic reviews, and critically appraised topics.
- Run a journal club so appraisal effort is shared across a department.
- Build evidence into standard workflows — protocols, order sets, and documentation templates.
- Keep a personal critically appraised topic file for questions that recur in your caseload.
- Contribute back. Program evaluation data and single-subject designs from routine practice are legitimate evidence and are how the profession's evidence base grows.
An occupational therapist reads a study of 3,000 clients reporting that a new hand exercise program produced a statistically significant improvement in grip strength (p = 0.001) with a mean gain of 1.4 pounds, where the established minimal clinically important difference is 12 pounds. How should the therapist interpret this finding?
An occupational therapist wants to determine whether a sensory-based intervention improves classroom participation for children with autism spectrum disorder compared with a standard classroom accommodation program. Which formulation BEST structures this as a searchable clinical question?
A therapist finds a well-conducted randomized controlled trial supporting an intensive six-hour-per-day intervention protocol for a client's diagnosis. The client meets every inclusion criterion but works full time and states clearly that they will not stop working. What does evidence-based practice require?