1.3 High Availability, Redundancy & Disaster Recovery

Key Takeaways

  • In a CUCM cluster, the Publisher holds the master read/write copy of the Informix Dynamic Server (IDS) database, whereas Subscribers maintain read-only local copies for real-time call processing and endpoint registration.

  • Database replication health is verified using 'utils dbreplication runtimestate', where replication status code 2 confirms that all replication channels and database tables are fully synchronized across the cluster.

  • Cisco Unified Survivable Remote Site Telephony (SRST) enables a branch Cisco IOS XE gateway to autonomously provide SIP registrar and call-routing capabilities when WAN connectivity to the centralized CUCM cluster is lost.

  • Enhanced SRST (E-SRST) provisions the branch router from CUCM data, so phone, line, speed-dial, and hunt-group settings survive in fallback without hand-built CLI; note that the v2.0 blueprint's high-availability objective excludes SRST.

  • The Disaster Recovery System (DRS) performs scheduled cryptographic backups of CUCM and Unity Connection components to remote SFTP storage, requiring the Publisher node to be restored and verified before restoring Subscribers.

Last updated: October 2026

1.3 High Availability, Redundancy & Disaster Recovery

Note

Blueprint objective 1.1.e (high availability) explicitly excludes SRST. SRST is kept in this section only as background for centralized designs; spend most of your time on cluster redundancy, CUCM Groups, keepalives, and disaster recovery.

Enterprise collaboration solutions must deliver continuous dial-tone availability, survivable branch telephony, and rapid disaster recovery capabilities. A comprehensive resilience strategy spans server role segregation, relational database replication, keepalive failure detection, remote branch survivability through Cisco IOS XE gateways, and reliable backup mechanisms.


1. CUCM Publisher-Subscriber Roles & Database Replication

Cisco Unified Communications Manager utilizes a clustered architecture based on the IBM Informix Dynamic Server (IDS) database engine.

+-----------------------------------------------------------------------------------+
|                         CUCM CLUSTER DATABASE ARCHITECTURE                        |
|                                                                                   |
|        PUBLISHER NODE (Primary)                     SUBSCRIBER 1 (Call Processing)|
|   +-------------------------------+           +-------------------------------+   |
|   | Informix Dynamic Server (IDS) |           | Informix Dynamic Server (IDS) |   |
|   | Master Database (Read/Write)  |           | Local Replicate (Read-Only)   |   |
|   +---------------+---------------+           +---------------+---------------+   |
|                   |                                           ^                   |
|                   |=== IDS Enterprise Replication (TCP 1500) =|                   |
|                   |                                           |                   |
|                   v                                           v                   |
|   +---------------+---------------+           +---------------+---------------+   |
|   | Administration & Provisioning |           | Endpoints & Trunks Register   |   |
|   | AXL, CCMAdmin, UDS, Extension |           | Call Control, Media, SIP, CTI |   |
|   +-------------------------------+           +-------------------------------+   |
+-----------------------------------------------------------------------------------+

Publisher and Subscriber Responsibilities

  • Publisher Node: Maintains the single master read/write copy of the Informix database. All administrative configurations, provisioning workflows, Cisco Unified Serviceability operations, and Administrative XML (AXL) API write requests must execute directly against the Publisher. Only one active Publisher exists per CUCM cluster.
  • Subscriber Nodes: Maintain local, read-only replicated copies of the Informix database. Subscribers handle real-time signaling, call processing, endpoint registration (SIP and SCCP), CTI routing, and media resource termination. If the Publisher node fails or becomes isolated, existing active calls continue uninterrupted, and Subscribers continue processing new calls and registrations using their local read-only database copies. However, administrative configuration modifications remain blocked until the Publisher recovers.

Database Replication Architecture & Verification

Database changes made on the Publisher are replicated across the cluster using IDS Enterprise Replication (ER) over TCP port 1500 (control) and TCP port 1501 (data).

Administrators verify replication status via the CUCM platform CLI using the command:

admin: utils dbreplication runtimestate

The output reports cluster connectivity and assigns an integer replication status code to each node:

Replication Status CodeStatus MeaningTechnical Description & Remediation Action
0InitializationReplication is being set up. Remaining in this state for more than an hour suggests a setup failure.
1Number of replicates incorrectSetup is still in progress (rarely seen in current releases). A long stay here also indicates a failure.
2GoodOptimal state. Logical connections are established and the tables match the other servers in the cluster.
3Mismatched tablesThe node is connected, but one or more tables are out of sync. Run utils dbreplication repair.
4Setup failed / droppedReplication setup did not succeed; investigate connectivity, DNS, and NTP, then usually reset replication.

Database Replication Troubleshooting Commands

  • utils dbreplication repair all: Compares table rows across all Subscribers against the Publisher and synchronizes mismatched tables without resetting the replication catalog.
  • utils dbreplication reset all: Completely tears down and rebuilds the Informix replication catalog, re-establishing replication queues across all cluster nodes.
  • utils dbreplication stop / utils dbreplication start: Halts and resumes the Informix replication service processes.

2. Keepalive Mechanisms & Failure Detection

Rapid failover depends on continuous keepalive heartbeats between endpoints, trunks, and call-processing subscribers.

SIP Trunk OPTIONS Ping

To monitor the operational status of SIP trunks connected to CUBE gateways, Cisco Unified Border Elements, or external PBXs, CUCM uses the SIP OPTIONS ping keepalive mechanism:

  • CUCM periodically transmits a SIP OPTIONS request to the destination IP address.
  • The remote device responds with 200 OK, 404 Not Found, or another SIP status code.
  • By default, the SIP Profile pings an in-service destination every 60 seconds and an out-of-service destination every 120 seconds, retransmitting an unanswered OPTIONS after the Ping Retry Timer (500 ms default). When a destination stops answering, CUCM marks it out of service and routes outbound calls to the next device in the route group without waiting for INVITE timers to expire.

Endpoint Registration Keepalives

  • SIP IP Phones: The phone's first REGISTER asks for 3600 seconds (Timer Register Expires in the SIP Profile), but CUCM answers with Expires: 120, taken from the SIP Station KeepAlive Interval service parameter. The phone therefore re-registers about every 115 seconds (120 minus the 5-second Timer Register Delta) and also sends keepalive registrations to its backup subscriber. If the primary stops answering or the TCP connection drops, the phone registers to the next subscriber in its CUCM Group.
  • SCCP IP Phones: Cisco Skinny Client Control Protocol endpoints exchange bidirectional StationKeepalive messages every 30 seconds. If a phone misses three consecutive keepalives (90 seconds total), it concludes the TCP connection has severed and initiates failover.
  • TFTP Redundancy: When booting, IP phones request configuration files (<SEP_MAC>.cnf.xml) from the TFTP server specified in DHCP Option 150. DHCP Option 150 supports an array of multiple IP addresses (e.g., Primary TFTP at Data Center 1, Secondary TFTP at Data Center 2), providing high-availability firmware and configuration loading.

3. Remote Site Telephony: Cisco IOS XE SRST & E-SRST

In centralized call processing models, a wide area network failure cuts off remote branch offices from the centralized CUCM cluster. Cisco Unified Survivable Remote Site Telephony (SRST) addresses this vulnerability by enabling the branch Cisco IOS XE router to assume local call control.

+-----------------------------------------------------------------------------------+
|                             SRST FAILOVER WORKFLOW                                |
|                                                                                   |
|  1. NORMAL OPERATION                                                              |
|     Branch IP Phone ======= IP WAN (SIP Signaling) =======> Central CUCM Cluster  |
|                                                                                   |
|  2. WAN FAILURE DETECTED                                                          |
|     Branch IP Phone ---X-- IP WAN (Keepalive Timeout) ----X- Central CUCM Cluster  |
|            |                                                                      |
|            v (Local Fallback Registration)                                        |
|     Cisco IOS XE Gateway (SRST / E-SRST Active)                                   |
|            |                                                                      |
|            v (PSTN Fallback Dialing)                                              |
|     Local PSTN Trunk (ISDN PRI / FXO / CUBE SIP)                                  |
+-----------------------------------------------------------------------------------+

SIP SRST vs. Enhanced SRST (E-SRST)

  • Standard SIP SRST: The IOS XE router acts as a lightweight SIP proxy and registrar. Phones register locally to the router using their pre-configured credentials. Basic dial plans, call transfers, and local PSTN breakout operate smoothly. However, user-specific features like speed dials, shared lines, hunt groups, and custom softkeys are unavailable unless manually configured via CLI.
  • Enhanced SRST (E-SRST): Removes most of the manual CLI work. The branch router's fallback configuration is provisioned from CUCM data, so directory numbers, line appearances, speed dials, hunt groups, and call-forward settings carry over. During a WAN failure, users see behavior close to normal CUCM operation.

Cisco IOS XE SIP SRST Configuration

The following verified configuration enables SIP SRST on a branch Cisco IOS XE router:

! Enable the SIP registrar so phones can register during fallback
voice service voip
 allow-connections sip to sip
 sip
  registrar server expires max 600 min 60
!
! Global SIP SRST limits
voice register global
 max-dn 100
 max-pool 50
!
! Phones from this subnet may register while in fallback
voice register pool 1
 id network 10.10.20.0 mask 255.255.255.0
 dtmf-relay rtp-nte
 codec g711ulaw
!
! Emergency calls leave through the local PRI
dial-peer voice 911 pots
 destination-pattern 911
 port 0/1/0:23
 forward-digits 3

Key configuration parameters:

  • registrar server: Turns on the router's SIP registrar. Without it, the router rejects the REGISTER requests that phones send during fallback.
  • max-pool and max-dn: Limit how many phones and directory numbers can register simultaneously.
  • voice register pool with id network: Defines which phones are allowed to register in fallback.
  • Phones learn where to fall back from the SRST Reference that CUCM assigns through the device pool (usually the router's LAN address); the router does not advertise itself.

4. Cisco Unified Communications Manager Groups

High availability at the endpoint layer is governed by Cisco Unified Communications Manager Groups (CUCM Groups).

  • A CUCM Group defines a prioritized list of up to three call-processing subscribers:
    1. Primary Subscriber: Handles active registration and call processing under normal conditions.
    2. Secondary Subscriber: Standby call processing agent that assumes registration upon primary failure.
    3. Tertiary Subscriber: Final fallback server if both primary and secondary nodes fail.
  • CUCM Groups are assigned to Device Pools, which are then associated with IP phones, voice gateways, and SIP trunks. By distributing different Device Pools across distinct primary and secondary subscribers, administrators balance call processing load across the cluster.

5. Disaster Recovery: DRS & System Restoration

When hardware failures, storage corruption, or catastrophic site events compromise a cluster, the Disaster Recovery System (DRS) orchestrates cluster-wide data restoration.

DRS Architecture & Backup Components

DRS is integrated into the CUCM platform and operates an agent-based architecture:

  • Master Agent (MA): Executes exclusively on the Publisher node, coordinating backup and restore schedules, authenticating storage credentials, and instructing cluster nodes.
  • Local Agent (LA): Executes on all cluster nodes (Publisher and Subscribers) to execute local file archiving and database export.
  • Storage Repository: Backups are written over SFTP (SSH File Transfer Protocol) to a dedicated, off-cluster storage server. Network File System (NFS) and Windows SMB shares are unsupported.

DRS backups are organized by feature:

  • UCM: The configuration database, platform settings, TFTP files (such as custom ring tones and backgrounds), and certificate stores of every node in the cluster.
  • CDR_CAR: The CDR repository and CDR Analysis and Reporting (CAR) data.
  • IM and Presence and Cisco Unity Connection have their own DRS features and are backed up from their own clusters.

Cluster Restoration Sequence

Restoring a CUCM cluster from DRS backups must adhere to a strict, non-negotiable sequence:

+-----------------------------------------------------------------------------------+
|                         DRS RESTORATION SEQUENCE                                  |
|                                                                                   |
|  STEP 1: Fresh Base Install                                                       |
|  Deploy fresh virtual machines matching original IP, Hostname, and Version.       |
|                                                                                   |
|  STEP 2: Restore Publisher First                                                  |
|  Launch DRS on Publisher -> Select SFTP Backup Tar -> Execute Restore.            |
|                                                                                   |
|  STEP 3: Verify Master Database Health                                            |
|  Reboot Publisher -> Check CLI: 'utils service list' (Database running).          |
|                                                                                   |
|  STEP 4: Restore Subscriber Nodes                                                 |
|  Launch DRS on Publisher -> Select Subscriber nodes -> Execute Restore.           |
|                                                                                   |
|  STEP 5: Audit Database Replication                                               |
|  Execute 'utils dbreplication runtimestate' -> Confirm Status 2 (Healthy).        |
+-----------------------------------------------------------------------------------+

Attempting to restore Subscribers prior to completing and verifying the Publisher restoration will cause severe database corruption, orphan database catalogs, and broken replication loops.

Loading diagram...
Branch Phone Keepalive Expiration and SRST Fallback Registration
Test Your Knowledge

A collaboration administrator executes the command 'utils dbreplication runtimestate' in the CUCM CLI after a planned maintenance window. The output shows a replication status of 2 for all subscriber nodes. What does this status indicate regarding the Informix database replication health?

A

Replication has failed due to mismatched Informix transactional logs.

B

The Publisher node is undergoing database repair and temporary read-only locking.

C

Replication has not started and is currently disabled across all nodes.

D

Replication is in a healthy, synchronized state across all cluster nodes.

Test Your Knowledge

During a WAN outage, a branch router configured with voice register global and voice register pool 1 rejects every fallback SIP REGISTER from its branch phones. Which missing configuration most likely causes this?

A

registrar server under voice service voip > sip

B

telephony-service in global configuration mode

C

session target ipv4:10.10.20.1 under a VoIP dial peer that points to the CUCM publisher

D

mode cme under voice register global

Sections you finish are checked off in the contents.