8.3 Supply Chain Risk Management
Key Takeaways
- External supply risk management follows identify, assess, mitigate, and monitor—not one-time contingency binders.
- Compare time-to-survive with time-to-recover; a TTS shorter than TTR is a structural gap requiring buffers, alternates, or faster recovery.
- Mitigations include safety stock and safety time, live dual sourcing, capacity reservations, contractual continuity clauses, and design flexibility.
- Monitoring uses OTD, lead-time trends, and risk-register refresh so planning parameters stay honest after conditions change.
- Resilience has a cost: CPIM favors matching buffer and dual-source intensity to item criticality rather than uniform policies.
8.3 Supply Chain Risk Management
Quick Answer: Supply chain risk management (SCRM) for external supply is a continuous cycle—identify, assess, mitigate, and monitor—applied to supplier, logistics, geopolitical, quality, and demand shocks. CPIM expects practical mitigations: inventory and capacity buffers, dual sourcing, contractual protections, and visibility—not hope that MRP will absorb every disruption.
Even excellent sourcing fails when a plant fire, port closure, cyber event, or quality escape removes supply. Domain V therefore tests whether planners treat risk as a managed parameter set, not a surprise after a stockout.
The SCRM Cycle
| Stage | Question | Typical outputs |
|---|---|---|
| Identify | What can go wrong in external supply? | Risk register by item/supplier/lane |
| Assess | How likely and how severe? | Risk scores; heat maps; financial impact |
| Mitigate | What reduces likelihood or impact? | Dual source, buffers, contracts, redesign |
| Monitor | Is the risk changing? | Scorecards, early-warning KPIs, audits |
Risk is not only catastrophic events. Chronic lead-time variance, single-source electronics, and financially weak suppliers are everyday Domain V risks.
Identifying External Supply Risks
Common risk categories for purchased supply:
- Supplier operational — capacity loss, quality escape, labor strike, IT outage
- Commercial / financial — insolvency, abrupt price spikes, contract disputes
- Logistics — carrier failure, port congestion, customs delays
- Geopolitical / regulatory — tariffs, export controls, sanctions, regional conflict
- Environmental / force majeure — natural disaster, pandemic restrictions
- Cyber / data — EDI/ASN failure, ransomware at supplier or 3PL
- Demand-side amplification — bullwhip causing supplier overload
| Risk example | Early indicator | Planning symptom |
|---|---|---|
| Supplier financial stress | Late filings, longer payment asks | Missed ASNs, broken promises |
| Port congestion | Transit-time creep | Inflated in-transit inventory, late receipts |
| Quality drift | Rising PPM / complaints | Scrap, short kits, reschedules |
| Capacity crunch | Longer quoted lead times | Growing past-due POs |
Supply Chain Mapping and Event Monitoring
You cannot manage a risk you cannot see. Supply chain mapping documents the network past tier 1 — plants, distribution centers, carriers, ports, and critically the tier-2 and tier-3 suppliers your direct suppliers depend on. Mapping exposes hidden concentration: three apparently independent tier-1 suppliers that all buy the same resin from one tier-2 plant are a single point of failure, and only a map reveals it.
A usable map records, for each node: what it supplies, sole/single/multi-source status, the qualification lead time to replace it, geographic exposure, and the revenue at risk behind it. Event monitoring then watches that map continuously — severe weather, port congestion, financial-distress signals, sanctions and export-control listings, labor actions — and raises alerts against mapped nodes instead of a generic news feed.
Mapping and monitoring feed risk tolerance decisions. A mapped node carrying little revenue at risk may simply be accepted; a mapped single-source node sitting behind 40 percent of revenue justifies funding a second qualified source.
Exam trap: mapping and monitoring are risk identification activities, not mitigation. Discovering a single point of failure changes no exposure until a mitigation — dual sourcing, a buffer, a contractual commitment — is actually funded and implemented.
Assessing Likelihood and Impact
Assessment combines probability and impact (service, cost, compliance, brand). A useful simplification for exam scenarios:
| Impact → / Likelihood ↓ | Low impact | High impact |
|---|---|---|
| Low likelihood | Accept & monitor | Contingency plans, light buffers |
| High likelihood | Process fixes, dual options | Aggressive mitigation (dual source + buffer + executive SRM) |
Also separate time-to-recover (TTR) and time-to-survive (TTS). If TTS (how long you can operate through inventory and alternates) is shorter than TTR (how long the supplier needs to restore supply), you have a structural gap that inventory policy or dual sourcing must close.
Scenario: Time-to-Survive Gap
A sole-sourced resin has 10 days of on-hand + inbound coverage (TTS ≈ 10 days). The only qualified supplier’s documented recovery after a major outage is 6 weeks (TTR ≈ 42 days). Assessment says the firm cannot survive the plausible disruption. Mitigation must extend TTS (more buffer, alternate formula) and/or shrink TTR (second qualified site, mold tooling at alternate), not merely raise the forecast.
Mitigation Levers
Inventory and Time Buffers
- Safety stock protects against quantity uncertainty
- Safety lead time / safety time protects against timing uncertainty
- Decoupling stock at strategic BOM levels absorbs upstream shocks
- Hedge inventory ahead of known events (port strikes, seasonal freeze)
Buffers are expensive. Apply them preferentially to high-impact, long-TTR exposures—not uniformly to every SKU.
Dual and Multiple Sourcing
Qualifying a second source reduces sole-source failure impact. Keep the secondary source “alive” with a meaningful volume share or periodic orders so tooling, process capability, and commercial terms stay current. A paper dual source that has not run in two years is often a single source in disguise.
Capacity Reservation and Long-Term Agreements
Contracts can reserve capacity, define allocation in shortage periods, and set collaboration cadences. Capacity reservations turn supplier bottlenecks into planned constraints rather than surprises.
Contractual Risk Allocation
Contracts do not eliminate physical risk, but they clarify remedies and incentives:
| Clause type | Purpose |
|---|---|
| Lead-time / OTD remedies | Service commitments and escalation |
| Quality warranties & PPAP | Defect liability and change control |
| Force majeure | Defines excusable delay—and what is not covered |
| Volume commitments / take-or-pay | Secures capacity; creates buyer obligation |
| Business continuity requirements | Auditable recovery expectations |
| Indemnity / liability caps | Financial exposure boundaries |
| Termination & transition assistance | Exit path if performance collapses |
Design and Process Flexibility
Alternate materials, universal components, and postponed differentiation reduce dependence on a single purchased part. Engineering change is a risk tool when sourcing options are exhausted.
Visibility and Collaboration
ASN compliance, shared forecasts, risk alerts, and multi-tier mapping (knowing your supplier’s suppliers for critical items) shorten detection time. You cannot mitigate what you discover only at the dock.
Monitoring and Continuous Review
Mitigations decay. Dual sources lose qualification; buffers become obsolete; suppliers’ financial health changes. Monitoring should include:
- Supplier scorecards (OTD, quality, responsiveness)
- Lead-time trend charts versus planning parameters
- Risk-register refresh after major events or S&OP cycles
- Periodic business-continuity tests for strategic suppliers
- Audit of contractual SLAs and claim patterns
| Monitor metric | Trigger action |
|---|---|
| OTD < target for 2 periods | Increase safety time; SRM escalation |
| Lead time +20% vs parameter | Update planning lead time; review buffer |
| Single-source A-item, no alternate | Launch dual-source or redesign project |
| In-transit variance rising | Change Incoterms/lane; add visibility tools |
Disruption Playbooks for Planners
When disruption hits, CPIM-aligned responses are sequenced:
- Confirm true supply position — on-hand, open PO, ASN, alternate sites
- Protect constrained supply — allocation rules to priority MPS items / customers
- Deploy buffers deliberately — do not ration blindly
- Activate alternates — second source, spot buy, substitute
- Reschedule demand and maintenance — honest ATP/CTP
- Capture lessons — update risk register and parameters
Scenario: Dual Source Saves the Schedule
An electronics manufacturer splits a microcontroller 70/30 between Supplier A (Asia) and Supplier B (Mexico). A regional lockdown idles Supplier A for five weeks. Because Supplier B is warm and contractually obligated to surge within rated capacity, the firm shifts mix to 40/60 temporarily, burns two weeks of safety stock on A-only variants, and keeps the MPS for the top revenue family intact. Without the live dual source, TTS would have been exceeded in under two weeks.
Integrating Risk into Planning Parameters
Risk management must change numbers the planner owns:
- Planning lead time reflects realistic, risk-adjusted expectations—not marketing promises
- Order policies and horizons consider recovery time
- Safety stock formulas incorporate supply variability, not only demand variability
- S&OP escalates strategic sole-source exposures as capacity risks
If risk stays in a binder while MRP uses fantasy lead times, Domain V has failed in practice.
FMEA: Ranking Risks Instead of Arguing About Them
The ECM names failure mode and effects analysis (FMEA) as a required risk tool. FMEA converts opinion into a ranked list. For each failure mode you rate three factors from 1 to 10 and multiply them into a risk priority number (RPN):
RPN = Severity × Occurrence × Detection
- Severity (S) — how badly the failure hurts the customer or the schedule
- Occurrence (O) — how often the failure mode is expected
- Detection (D) — how hard it is to catch before it does damage. This is the rating candidates reverse: 10 means undetectable, 1 means it is caught immediately. Better detection lowers the score.
RPN therefore ranges from 1 to 1,000.
| Failure mode | S | O | D | RPN | Action |
|---|---|---|---|---|---|
| Sole-source resin plant outage | 9 | 3 | 8 | 216 | Qualify second source |
| Inbound customs delay | 5 | 6 | 3 | 90 | Add safety time to planning lead time |
| Wrong item master lead time | 6 | 4 | 9 | 216 | Add data audit — detection is the weak link |
| Carrier damages pallet | 4 | 5 | 2 | 40 | Accept; monitor claims |
Read the third and first rows together: identical RPN, completely different fixes. The resin risk needs a second source (attack occurrence); the bad master data needs an audit or system control (attack detection). Attacking the wrong factor is the classic FMEA exam trap. A high-severity item usually deserves attention even at a modest RPN, because severity is the one factor you cannot design away by inspecting harder.
Risk Management Standards and Guidance
The ECM expects recognition of the published frameworks, not memorized clause numbers:
- ISO 31000 — risk management principles, framework, and process. It is guidance, not a certifiable requirements standard, so an organization aligns with it rather than being audited to it.
- ISO 22301 — business continuity management systems; this one is certifiable and is where the business continuity plan (BCP) discipline formally lives.
- ISO 28000 — security management systems for the supply chain.
- ISO 9001 and ISO 14001 — quality and environmental management systems, which carry risk-based thinking into supplier qualification.
Security Requirements and Compliance
Security is a named ECM obligation covering both physical and cyber exposure, and both land on the planner.
| Domain | Representative controls | Planning consequence if it fails |
|---|---|---|
| Physical | Site access control, cargo seals, segregation of duties on inventory transactions, secure storage for high-value and controlled items | Shrinkage destroys record accuracy, so MRP plans against fiction |
| Cyber | Access rights in the ERP, master-data change control, secure EDI and supplier portals, ransomware recovery plans | A locked planning system stops releases, ATP promises, and receipts at once |
| Trade/customs | Programs such as the Customs-Trade Partnership Against Terrorism (C-TPAT) and equivalent authorized-operator schemes | Non-participation adds border inspection time, which is really added lead-time variability |
The planning insight is that a cyber incident is a capacity and lead-time event, not just an IT event. If the ERP is down for four days, you have lost four days of releases and receipts; treat recovery time as you would a supplier outage, and keep an offline picture of the top constrained items.
Key Trade-Offs the Exam Loves
| Choice | Upside | Downside |
|---|---|---|
| More safety stock | Higher TTS | Higher carrying cost / obsolescence |
| Dual sourcing | Continuity | Complexity, possible unit-cost rise |
| Tight contracts | Clear remedies | Negotiation time; rigid terms |
| Single global low-cost source | Unit price | Concentrated geopolitical/logistics risk |
| Local second source | Faster recovery | Higher price; capacity limits |
There is no free resilience. CPIM rewards selecting the mix of buffers, dual sourcing, and contracts that matches item criticality and TTR/TTS reality—and then monitoring so the mix stays valid.
What is the correct sequence for managing external supply risk in the SCRM cycle emphasized for CPIM-style planning?
A sole-sourced component has 12 days of cover including inbound (time-to-survive) but the supplier’s documented recovery after a major outage is 40 days (time-to-recover). What does this assessment imply?
Which mitigation best addresses sole-source disruption risk while keeping an alternate capable over time?
Which contractual element most directly supports continuity planning with a strategic supplier?