1.2 High Availability Architecture: Device Redundancy, FHRPs, and Stateful Switchover (SSO)
Key Takeaways
Stateful Switchover (SSO) continuously synchronizes protocol and forwarding state to a standby supervisor, so a switchover takes seconds or less without dropping Layer 2 links.
Nonstop Forwarding (NSF) with Graceful Restart allows the data plane to forward packets uninterrupted using existing CEF tables while routing protocols re-establish neighbor adjacencies.
Cisco StackWise-Virtual (SVL) combines dual physical modular switches into a single logical control plane across 10G/40G/100G Ethernet links, replacing Spanning Tree with Multichassis EtherChannel.
Dual-Active Detection (DAD) uses Fast Hello or ePAgP to detect a split-brain condition after all StackWise Virtual Links fail, and it forces the original active switch into recovery mode.
First Hop Redundancy Protocols (HSRP, VRRP, GLBP) protect client default gateways, and integrating Bidirectional Forwarding Detection (BFD) achieves sub-second failure detection.
High Availability Architecture: Device Redundancy, FHRPs, and Stateful Switchover (SSO)
High availability (HA) in modern enterprise networks ensures uninterrupted business operations by eliminating single points of failure. HA architectures combine physical hardware redundancy, resilient supervisor switchover engines, control plane non-disruptive restart protocols, multi-chassis virtualization, and rapid first-hop gateway failover.
Hardware Redundancy and Power Architecture
Resilient hardware design protects physical networking chassis from component failures:
- Power Supply Redundancy Modes:
- Redundant Mode (1+1 or N+1): One power supply unit (PSU) actively powers the chassis while an auxiliary PSU remains on standby. If the primary PSU fails, the secondary unit assumes 100% of the electrical load without interrupting chassis operation.
- Combined / Sharing Mode: Multiple PSUs operate concurrently to aggregate their total wattage, supporting dense Power over Ethernet (PoE+ and UPoE) line card configurations. However, if a single supply fails in combined mode, the switch may be forced to power down non-critical PoE ports.
- Grid Redundancy: Modular switches connect dual internal power supplies to separate AC electrical circuits (Feed A and Feed B) powered by independent Uninterruptible Power Supply (UPS) systems and backup diesel generators.
- Fan Tray and Module Redundancy: Chassis incorporate dual hot-swappable fan trays with variable-speed fans, redundant clock modules, and hot-swappable line cards that can be inserted or removed while the chassis is fully powered.
Supervisor Engine Redundancy Modes
Modular enterprise switches (such as Cisco Catalyst 9400 and 9600 series) host dual Supervisor Engines (Route Processors) inside the same physical chassis. The active supervisor manages the control and data planes, while the standby supervisor monitors the primary unit. The synchronization level between supervisors defines the redundancy mode.
+-------------------------------------------------------------+
| ACTIVE SUPERVISOR |
| Control Plane: OSPF, BGP, STP, LACP | Data Plane: CEF / FIB |
+-------------------------------------------------------------+
||
State Synchronization (IPC Mirroring)
||
+-------------------------------------------------------------+
| STANDBY SUPERVISOR |
| SSO Mode: Control State Mirrored | Hardware FIB Pre-loaded |
+-------------------------------------------------------------+
1. Route Processor Redundancy (RPR)
In RPR mode, the standby supervisor is only partially booted. Its operating system kernel is loaded into memory, but the control plane software is not initialized, and the routing tables are empty.
- Failover Behavior: When the active supervisor fails, the standby supervisor must complete its boot sequence, reset all line cards, re-initialize hardware ASICs, and restart routing protocols.
- Downtime: Failover takes 2 to 4 minutes, causing extensive packet loss and link reconvergence across the network.
2. Route Processor Redundancy Plus (RPR+)
In RPR+ mode, the secondary supervisor is fully booted and its operating system is running. Line card drivers are pre-initialized.
- Failover Behavior: Upon active supervisor failure, the standby takes control without resetting line cards, preserving physical Layer 1 link carrier states. However, the control plane protocols (STP, OSPF, EIGRP, BGP) must restart from scratch, and forwarding tables must be rebuilt.
- Downtime: Failover requires 30 to 60 seconds, preserving link states but dropping packets during protocol recalculation.
3. Stateful Switchover (SSO)
Stateful Switchover represents the gold standard for supervisor high availability. In SSO mode, the active supervisor continuously synchronizes hardware configuration, Layer 2 protocol states, interface status, and chassis environment data to the standby supervisor in real time via Inter-Process Communication (IPC).
- Data Plane Continuity: Cisco Express Forwarding (CEF) and Forwarding Information Base (FIB) tables are pre-populated on the standby supervisor's hardware ASICs before any failure occurs.
- Layer 2 Protocol Transparency: Spanning Tree Protocol (STP), Port Aggregation Control Protocol (PAgP), Link Aggregation Control Protocol (LACP), and 802.1X sessions are mirrored seamlessly. Switch ports remain in the forwarding state without link flaps.
- Downtime: Cisco Catalyst documentation has quoted SSO switchover in the 0 to 3 second range, compared with 30 to 60 seconds for RPR+ and minutes for RPR. Links stay up and established user sessions normally survive.
Supervisor Redundancy Modes Comparison
| Feature / Metric | Route Processor Redundancy (RPR) | Route Processor Redundancy Plus (RPR+) | Stateful Switchover (SSO) |
|---|---|---|---|
| Standby OS State | Partially booted (kernel loaded) | Fully booted and running | Fully booted, running, state-synced |
| Line Card Reset | Yes; all line cards are power-cycled | No; line cards maintain link state | No; line cards remain active |
| Layer 2 Protocols | Restarted from scratch; links flap | Restarted from scratch | Preserved seamlessly via IPC |
| CEF / FIB State | Completely flushed and rebuilt | Flushed and rebuilt | Pre-programmed into standby ASICs |
| Failover Duration | 2 to 4 minutes | 30 to 60 seconds | About 0 to 3 seconds (platform-dependent) |
| Traffic Impact | Total traffic outage | Outage during routing rebuild | Minimal loss; links stay up |
Nonstop Forwarding (NSF) with Graceful Restart
While SSO preserves Layer 2 connectivity, Layer 3 routing engines (OSPF, BGP, EIGRP, IS-IS) run as dynamic control plane processes. Immediately following an SSO switchover, the newly active supervisor lacks dynamic neighbor adjacency state.
Without assistance, neighboring routers would detect missed hello packets, tear down neighbor adjacencies, flush routing tables, and flood route withdrawal messages throughout the enterprise.
Nonstop Forwarding (NSF) works in tandem with SSO and IETF Graceful Restart (RFC 3623 for OSPF, RFC 4724 for BGP) to eliminate this problem:
- Continuous Forwarding: The newly active supervisor instructs the hardware switching ASICs to continue forwarding transit traffic using the existing CEF/FIB table programmed prior to switchover.
- Helper Adjacency Protection: The recovering router transmits a Graceful Restart signal to its neighbors (operating as NSF helper routers).
- Suppression of Route Flaps: The helper routers maintain active routing adjacencies and continue directing traffic toward the recovering router without advertising topology changes to the rest of the campus.
- Database Resynchronization: The recovering router re-learns neighbor states, rebuilds its Routing Information Base (RIB), refreshes its CEF table, and exits Graceful Restart mode with little or no packet loss.
Multi-Chassis Stacking: StackWise vs. StackWise-Virtual
Physical switch clustering unifies multiple chassis into a single logical management and forwarding domain:
Cisco StackWise and StackWise-480
- Target Platforms: Fixed-configuration access switches (e.g., Catalyst 9200, 9300).
- Physical Interconnect: Dedicated proprietary stacking cables forming a bidirectional, closed-loop ring with backplane bandwidths up to 480 Gbps or 1 Tbps.
- Control Architecture: Elects one Active switch, one Standby switch (synchronized via SSO), and multiple Member switches. All data forwarding is distributed across the ring.
Cisco StackWise-Virtual (SVL)
- Target Platforms: Modular and high-density distribution and core switches (Catalyst 9500, 9600 series), superseding legacy Virtual Switching System (VSS).
- Physical Interconnect: Eliminates specialized stacking cables by using standard 10G, 40G, or 100G Ethernet optical transceivers over StackWise-Virtual Links (SVL).
- Elimination of STP: Unifies two physical chassis into a single logical Layer 2/Layer 3 switch. Downstream access switches connect using Multichassis EtherChannel (MEC/LACP) distributed across both chassis, eliminating blocked Spanning Tree uplinks.
+-----------------------+ +-----------------------+
| Chassis 1 (Active) |=====| Chassis 2 (Standby) |
+-----------------------+ SVL +-----------------------+
\ \ / /
\ \ Dual-Active / /
\ \ Detection(DAD)/ /
\ +-------------+ /
\ /
Multichassis EtherChannel (MEC / LACP)
\ /
+-----------------------------+
| Access Layer Switch |
+-----------------------------+
Dual-Active Detection (DAD)
If all SVL links between the two chassis fail while both units remain powered, a split-brain condition occurs. Both switches assume the other has failed, and both attempt to operate as the Active switch, originating identical IP addresses, virtual MAC addresses, and routing IDs.
Dual-Active Detection (DAD) identifies this condition using secondary out-of-band communication paths:
- DAD Mechanisms: StackWise Virtual supports two methods: Fast Hello over up to four dedicated Layer 2 links between the switches, and enhanced PAgP (ePAgP) carried over a downstream MEC to a neighbor that supports it.
- Recovery Action: When the last SVL fails, the standby switch cannot tell whether the active switch died, so it takes over the active role. When DAD then detects that both switches are active, the original active switch enters Recovery Mode and shuts down all its interfaces except the SVL and management interfaces. After the SVL is repaired, the recovering switch reloads and rejoins as the standby.
First Hop Redundancy Protocols (FHRP) and Convergence
First Hop Redundancy Protocols ensure client endpoints maintain continuous default gateway access if an upstream router fails.
- HSRP (Hot Standby Router Protocol): Cisco proprietary. Active and Standby routers share a virtual IP and virtual MAC (
0000.0c07.acXXfor v1;0000.0c9f.fXXXfor v2, where XXX is the hex group number). Default hello interval is 3 seconds; hold time is 10 seconds. - VRRP (Virtual Router Redundancy Protocol): Open IETF standard (RFC 5798). Master and Backup routers share a virtual IP and virtual MAC (
0000.5e00.01XXfor IPv4;0000.5e00.02XXfor IPv6). Default advertisement interval is 1 second; dead interval is 3 seconds. - GLBP (Gateway Load Balancing Protocol): Cisco proprietary. An Active Virtual Gateway (AVG) responds to client ARP requests with distinct virtual MAC addresses assigned to up to four Active Virtual Forwarders (AVFs), providing active-active egress load sharing across gateways.
Sub-Second Failover with BFD
Default FHRP hello timers (1 to 3 seconds) result in a 3- to 10-second failover window—unacceptable for voice, video, or financial transactions. While timers can be manually tuned to sub-second values (e.g., 200ms hello / 600ms hold), doing so across hundreds of VLANs places severe processing burdens on switch CPUs.
Bidirectional Forwarding Detection (BFD) resolves this:
- BFD sends lightweight keepalives at short intervals (for example, 50 ms with a multiplier of 3); many platforms offload BFD to hardware or line-card software.
- When BFD detects a link disruption, it immediately notifies HSRP, VRRP, or routing protocols.
- Total failover convergence drops to sub-second intervals (100–200 ms) without placing excessive load on the control plane CPU.
Failover Convergence Timers Comparison
| Redundancy Mechanism | Detection Interval | Total Failover Window | Forwarding Plane Impact |
|---|---|---|---|
| Default HSRP (v1/v2) | 3s Hello / 10s Hold | 10 seconds | Significant dropped packets; TCP stall |
| Default VRRP (v2/v3) | 1s Advert / 3s Dead | 3 seconds | Dropped VoIP calls, session stalls |
| Tuned FHRP Timers | 200ms Hello / 600ms Hold | 600 milliseconds | Minimal packet loss; high CPU usage |
| FHRP with BFD Integration | 50ms Echo / 150ms Detect | 150 milliseconds | Sub-second failover; low CPU overhead |
| SSO Supervisor Switchover | Supervisor failure detection | About 0 to 3 seconds | Links stay up; minimal Layer 2 loss |
| SSO with NSF Graceful Restart | Forwarding continues on existing CEF/FIB | Routing re-converges in the background | Minimal loss while neighbors stay up |
What is the key operational distinction between Route Processor Redundancy Plus (RPR+) and Stateful Switchover (SSO)?
RPR+ powers down the standby supervisor until a hardware fault occurs, whereas SSO keeps the standby supervisor booted
RPR+ supports active-active forwarding across both supervisors at once, while SSO always operates in an active-standby mode
SSO keeps protocol state and forwarding tables synchronized so links do not reset, whereas RPR+ must restart its protocols
SSO requires dedicated external stacking cables, whereas RPR+ operates over standard Ethernet links
During a supervisor switchover on a router configured with Nonstop Forwarding (NSF) and Graceful Restart, what action is performed by an adjacent NSF helper router?
The helper router keeps the adjacency and keeps forwarding to the restarting router without a topology flap
The helper router immediately flushes all prefixes learned from the recovering router to prevent routing loops
The helper router initiates a Spanning Tree topology change notification across the campus backbone
The helper router takes over the active supervisor role on the recovering switch chassis
In a Cisco StackWise Virtual pair, all StackWise Virtual Links (SVLs) fail while both switches stay powered, and Dual-Active Detection confirms that both are now active. What happens next?
Both switches keep the active role and load-share traffic using the same virtual MAC address
The original active switch powers off both of its power supplies to avoid a split-brain condition
The former standby switch reboots every 30 seconds until an SVL comes back up
The original active switch enters Recovery Mode and shuts down all interfaces except the SVL and management interfaces
Sections you finish are checked off in the contents.