1.1 Collaboration Architecture & Deployment Models

Key Takeaways

  • Cisco Unified Communications Manager (CUCM) supports four foundational deployment models: Single-site, Multi-site with centralized call processing, Multi-site with distributed call processing, and Clustering over the WAN.

  • Clustering over the WAN enforces an absolute maximum round-trip time (RTT) of 80 ms between any two cluster nodes, paired with dedicated QoS bandwidth reserved for Intra-Cluster Communication Signaling (ICCS).

  • Centralized call processing deploys call control at a primary headquarters or data center while using Call Admission Control (CAC) and Survivable Remote Site Telephony (SRST) to maintain branch survivability during WAN outages.

  • Cisco Webex Dedicated Instance provisions private cloud-hosted CUCM, Unity Connection, and Expressway virtual nodes, enabling enterprises to retain custom dial plans and hardware endpoints without managing on-premises server infrastructure.

  • In subscriber sizing and redundancy, the 1:1 server redundancy model pairs each primary subscriber with a dedicated backup subscriber to guarantee zero-loss registration failover, whereas the 2:1 model pairs two primary subscribers with one shared backup subscriber.

Last updated: October 2026

1.1 Collaboration Architecture & Deployment Models

Enterprise collaboration systems demand robust, scalable, and resilient architectures capable of handling real-time voice, video, messaging, and presence services. Designing an enterprise collaboration infrastructure requires selecting the optimal deployment model based on geographic dispersion, network topology, round-trip latency, available bandwidth, and organizational governance requirements.


1. Collaboration Architecture Overview

Modern collaboration solutions span three primary deployment models:

  1. On-Premises Collaboration: All call-processing infrastructure, messaging servers, session border controllers, and media resources reside within private customer-managed data centers on dedicated computing platforms, such as Cisco Unified Computing System (UCS) servers running VMware ESXi. The enterprise exercises complete authority over data sovereignty, media routing, cryptographic key management, dial plan logic, and software maintenance lifecycles.
  2. Cloud Collaboration (Webex Calling): Multi-tenant cloud-native call control operated and maintained by Cisco. Call routing, feature delivery, and administrative management occur via the Webex Control Hub. Organizations transition from capital expenditures (CapEx) to predictable operational expenditures (OpEx), leveraging continuous feature updates without managing hypervisors, server hardware, or operating system patches.
  3. Hybrid Collaboration: Bridges on-premises platforms (such as Cisco Unified Communications Manager) with Webex cloud services. Common topologies include pairing on-premises call processing with Webex Meetings, Webex App messaging, Webex Video Mesh nodes for localized video switching, or utilizing Cisco Unified Border Element (CUBE) as a Local Gateway (LGW) to connect on-premises Public Switched Telephone Network (PSTN) circuits to Webex Calling.

Architectural Model Comparison

Design AttributeOn-Premises ArchitectureHybrid ArchitectureCloud Architecture (Webex Calling)
Control & Governance100% customer-controlled infrastructure and dial plansSplit control; on-prem call control with cloud servicesCentrally managed by Cisco via Webex Control Hub
Data SovereigntyAbsolute; media, signaling, and CDR records remain on-premMixed; sensitive call legs local, messaging in cloudRegulated cloud storage across regional data centers
Cost ModelHigh initial CapEx (hardware, perpetual licenses, SAN)Balanced CapEx / OpEx transition modelPredominantly OpEx per-user monthly/annual subscription
Maintenance OverheadHigh; manual OS patching, ESXi upgrades, SU maintenanceModerate; on-premises appliances require local careMinimal; zero client hypervisor or OS maintenance
PSTN FlexibilityLocal SIP trunks, T1/E1 PRIs, FXO/FXS analog gatewaysCUBE Local Gateway (LGW), on-prem carrier integrationCloud-Connected PSTN (CC-PSTN) or Local Gateway

2. Cisco Unified Communications Manager Deployment Models

Cisco Unified Communications Manager (CUCM) supports four established architectural deployment models defined by physical location, call control distribution, and IP WAN dependency.

Single-Site Deployment Model

In a single-site model, all call-processing nodes (Publisher and Subscribers), voicemail servers (Cisco Unity Connection), applications, media resources, and endpoints reside at a single physical campus or data center.

  • Signaling and Media Flow: Signaling traverses the local high-speed Local Area Network (LAN) directly to the CUCM cluster. Real-time Transport Protocol (RTP) media streams flow directly between endpoints or between endpoints and local gateways.
  • Bandwidth & Codecs: LAN bandwidth is abundant (typically 1 Gbps to 10 Gbps access layers), allowing widespread use of uncompressed, high-fidelity wideband and narrowband codecs such as G.711u/a and G.722 without WAN Call Admission Control (CAC) constraints.
  • PSTN Interface: Local T1/E1 PRI circuits or local SIP trunks terminate on local voice gateways.

Multi-Site Deployment with Centralized Call Processing

The centralized call processing model places a single CUCM cluster at a central headquarters or data center site, serving both local endpoints and geographically dispersed branch offices across an IP WAN.

  • WAN Dependence: Signaling for all branch endpoints (SIP registration, call setup, Supplementary Services) traverses the IP WAN to reach the centralized CUCM cluster.
  • Call Admission Control (CAC): WAN links possess constrained bandwidth. Administrators must enforce Locations CAC to prevent voice and video calls from oversubscribing WAN links and causing packet drop, jitter, and delay.
  • Survivability (SRST): If the IP WAN connection between a branch and the centralized CUCM cluster drops, branch phones lose connectivity to their call processing agent. To maintain emergency dialing and local site calling, branch routers must run Cisco Unified Survivable Remote Site Telephony (SRST), acting as temporary call control agents.
  • Bandwidth Conservation: Low-bitrate codecs (such as G.729 or Opus narrowband) are enforced across the IP WAN using Region settings to preserve circuit capacity.

Multi-Site Deployment with Distributed Call Processing

The distributed call processing model consists of multiple independent, autonomous CUCM clusters or call-processing entities located at different regional campuses or data centers.

  • Intercluster Communication: Clusters interconnect across the IP WAN with SIP intercluster trunks, often aggregated through a Cisco Unified CM Session Management Edition (SME) cluster. The Intercluster Lookup Service (ILS) with Global Dial Plan Replication can advertise directory URIs and number patterns between clusters so that each cluster does not need hand-built routes to every other cluster.
  • Operational Autonomy: Each campus manages its own local dial plan, local PSTN connections, and local survivability. A WAN outage severs inter-site calling but leaves intra-site calling and external PSTN access completely uninterrupted.
  • Scalability: Overcomes the ceiling of a single cluster. Cisco's Release 15 sizing guide allows up to 50,000 registered phones in a standard cluster (four primary call-processing subscribers on the Large OVA at 12,500 endpoints each), so global enterprises with more devices, or with strong regional autonomy needs, deploy several clusters.

Clustering over the WAN

Clustering over the WAN is a specialized deployment model where nodes belonging to the same logical CUCM cluster are distributed across distinct physical sites separated by an IP WAN or metropolitan area network (MAN).

+-----------------------------------------------------------------------------------+
|                             SINGLE LOGICAL CUCM CLUSTER                           |
|                                                                                   |
|        DATA CENTER A (PRIMARY)                     DATA CENTER B (SECONDARY)      |
|   +-------------------------------+           +-------------------------------+   |
|   |  CUCM Publisher (IDS Read/Write)|         |  CUCM Subscriber 2 (Call Proc)|   |
|   |  CUCM Subscriber 1 (Call Proc)|           |  CUCM TFTP / Backup Node      |   |
|   +---------------+---------------+           +---------------+---------------+   |
|                   |                                           |                   |
|                   +===================+=======================+                   |
|                                       |                                           |
|                        DEDICATED METRO / WAN IP FABRIC                            |
|                        Latency: <= 80 ms Round-Trip Time (RTT)                    |
|                        Bandwidth: Dedicated QoS for ICCS Signaling                |
+---------------------------------------+-------------------------------------------+
                                        |
                   +--------------------+--------------------+
                   |                                         |
           BRANCH 1 (CENTRALIZED)                    BRANCH 2 (SRST FALLBACK)
     +-----------------------------+           +-----------------------------+
     | IP Phones -> Sub 1 (Primary)|           | IP Phones -> Sub 2 (Primary)| 
     | Cisco ISR Gateway (PSTN)    |           | Cisco ISR Gateway (SRST)    |
     +-----------------------------+           +-----------------------------+

3. Clustering over the WAN: Engineering Specifications

Deploying a single CUCM cluster across geographic boundaries requires strict adherence to engineering latency, bandwidth, and redundancy guidelines.

Latency Boundaries

  • Maximum Round-Trip Time (RTT): The absolute maximum allowable network delay between any two nodes in a CUCM cluster over the WAN is 80 ms RTT (equivalent to 40 ms one-way delay).
  • Consequences of Excessive Latency: Latency exceeding 80 ms degrades Informix Dynamic Server (IDS) database replication, induces replication timeouts, causes out-of-order Intra-Cluster Communication Signaling (ICCS), and produces race conditions during call setup and supplementary service invocation.

Bandwidth Allocation for ICCS

Intra-Cluster Communication Signaling (ICCS) handles inter-node state exchange, including station registration events, call state notifications, and feature coordination (such as shared lines, directed call park, and BLF presence).

ICCS runs over TCP port 8002 (TCP 8003 carries intracluster traffic to the CTIManager). Because ICCS traffic directly affects call processing and dial tone delivery, it must be prioritized in the network QoS framework using Class-Based Weighted Fair Queuing (CBWFQ) or Priority Queuing (marked as CS3 or AF31).

Cisco's design guidance gives two separate bandwidth components for clustering over the WAN:

  • ICCS: at least 1.544 Mbps (T1) for every 10,000 busy hour call attempts (BHCA) between the sites, when directory numbers are not shared across sites. For more traffic, Cisco's guideline formula is:

ICCS bandwidth (Mbps)=Total BHCA10,000×(1+0.006×RTT in ms)\text{ICCS bandwidth (Mbps)} = \frac{\text{Total BHCA}}{10{,}000} \times (1 + 0.006 \times \text{RTT in ms})

  • Database and other inter-server traffic: an additional 1.544 Mbps for every subscriber that is remote from the publisher.

Worked example: 15,000 BHCA between two sites with an 80 ms RTT, and three subscribers remote from the publisher:

ICCS=1.5×(1+0.48)=2.22 Mbps\text{ICCS} = 1.5 \times (1 + 0.48) = 2.22 \text{ Mbps}

Database=3×1.544=4.632 Mbps\text{Database} = 3 \times 1.544 = 4.632 \text{ Mbps}

The sites need roughly 6.85 Mbps (about 7 Mbps) reserved for cluster traffic, before any voice or video media is added. Shared directory numbers across the sites require extra ICCS bandwidth.

Database Replication vs. Call Processing Failover

It is vital to distinguish between database synchronization traffic and call-processing failover:

  • Database Replication: Transmitted via the Informix Dynamic Server (IDS) enterprise engine over TCP ports 1500 and 1501. The Publisher pushes data modifications to Subscribers asynchronously. While a WAN disruption halts administrative updates across the WAN, active Subscribers continue processing local calls using their read-only local database copies.
  • Call Processing Failover: Governed by endpoint signaling keepalives. When an active Subscriber becomes unreachable, IP endpoints detect the missing keepalives and fail over to their configured secondary Subscriber in their assigned Cisco Unified Communications Manager Group.

Server Redundancy Models (1:1 vs. 2:1)

A standard CUCM cluster has at most four primary and four backup call-processing subscribers (a megacluster allows up to eight pairs), and the publisher, TFTP, music on hold and other media nodes all count toward a maximum of 21 servers. Each Large-OVA subscriber can register up to 12,500 endpoints (Medium 10,000; Small 3,750). Two redundancy models are used:

  1. 1:1 Server Redundancy Model:

    • Each primary call-processing subscriber is paired with a dedicated secondary standby subscriber (e.g., Sub1 paired with Sub2; Sub3 paired with Sub4).
    • Under normal operations, 50% of the nodes process calls while 50% remain idle or handle minimal baseline load.
    • Advantage: Provides 100% capacity absorption during a node failure. If Sub1 crashes, all of its registered phones (up to 12,500 on the Large OVA) migrate to Sub2 without exceeding registration ceilings or CPU capacity.
    • Disadvantage: Requires doubling the virtual machine footprint and hypervisor server resources.
  2. 2:1 Server Redundancy Model:

    • Two primary call-processing subscribers share a single dedicated backup subscriber (e.g., Sub1 and Sub2 both back up to Sub3).
    • Advantage: Reduces virtual machine footprint, vCPU consumption, and compute hardware overhead by 25% to 33%.
    • Engineering Rule: The combined total of active registrations across Sub1 and Sub2 must never exceed the capacity of the shared backup node (12,500 endpoints for the Large OVA). If both Sub1 and Sub2 fail simultaneously, Sub3 will reject registration attempts that exceed its rated capacity.

4. Cloud Integration: Webex Calling & Dedicated Instance

Enterprises transitioning toward cloud architectures can integrate on-premises environments with Cisco cloud collaboration platforms:

Webex Dedicated Instance (DI)

Webex Dedicated Instance provides dedicated, private CUCM, Cisco Unity Connection, Cisco Unified IM and Presence, and Cisco Expressway virtual machines running inside Cisco Webex cloud data centers.

  • Architecture: Unlike multi-tenant Webex Calling, each customer receives dedicated virtual instances identical to on-premises software releases.
  • Connectivity: Connects to the enterprise corporate network via Webex Edge Connect (dedicated, SLA-backed private Layer 3 peering via Equinix Fabric or partner interconnect) or private IPsec VPN tunnels.
  • Use Case: Ideal for large enterprises with complex dial plans, extensive third-party Computer Telephony Integration (CTI) recording platforms, analog integrations, or specialized Cisco IP phone models that cannot operate on multi-tenant cloud platforms.

Hybrid Calling via Local Gateway (CUBE)

For organizations utilizing multi-tenant Webex Calling while retaining on-premises PSTN connectivity or legacy analog paging systems, a Cisco IOS XE router running CUBE is configured as a Local Gateway (LGW).

  • A registration-based Local Gateway registers to Webex Calling over SIP TLS (destination TCP port 8934); a certificate-based Local Gateway uses a mutual TLS trunk (destination TCP port 5062). Section 9.6 covers both.
  • Calls to and from the cloud are normalized, translated, and routed to local on-premises PRI or SIP circuits, preserving local carrier contracts and emergency dispatch routing.

5. Realistic Exam Scenario: Enterprise WAN Architecture

Scenario Description

An enterprise with corporate headquarters in New York (Site A) operates three remote locations:

  • Site B (Newark Data Center): 8 miles away; fiber link with 3 ms RTT.
  • Site C (London Regional Office): Intercontinental MPLS link with 68 ms RTT and 0.5% packet loss.
  • Site D (Tokyo Branch Office): Transpacific link with 145 ms RTT and 1.2% packet loss.

The enterprise requirements state:

  1. Site A and Site B must operate with instantaneous zero-loss failover for 15,000 phones.
  2. The architecture must minimize server overhead where technically feasible.
  3. The enterprise wishes to maintain unified administration across sites whenever engineering limits permit.

Architectural Decision Analysis

  • Site A to Site B Design: With 3 ms RTT, Site A and Site B easily fall under the 80 ms boundary. They should deploy Clustering over the WAN with a 1:1 redundancy model (Subscribers in New York backed up by Subscribers in Newark). This satisfies the requirement for zero-loss line-rate failover for 15,000 endpoints.
  • Site C (London) Design: With 68 ms RTT, London is within the 80 ms RTT limit for Clustering over the WAN. However, because intercontinental bandwidth is expensive, the architect must evaluate whether to place a Subscriber node in London with dedicated 1.544 Mbps ICCS bandwidth or to implement Centralized Call Processing with SRST. If local survivability and local call processing independence are preferred without dedicating continuous ICCS bandwidth, Centralized Call Processing with local SRST on an ISR 4000 gateway is the recommended design.
  • Site D (Tokyo) Design: With 145 ms RTT, Tokyo violates the 80 ms RTT maximum between cluster servers, so the architect must not place a cluster node in Tokyo. The 80 ms limit applies only between servers in the cluster: Tokyo phones could still register to New York in a centralized design, but the 145 ms path and 1.2% loss would hurt signaling responsiveness and media quality. A local cluster (distributed call processing) with an intercluster SIP trunk to New York, or a cloud option such as Webex Calling with a local PSTN connection, gives Tokyo local call control.
Loading diagram...
Enterprise Multi-Site Collaboration Topology with Latency Boundaries
Test Your Knowledge

An enterprise is designing a multi-site CUCM cluster across two data centers separated by a metropolitan optical network. Network telemetry reports a round-trip time (RTT) of 62 ms between the data centers. Which statement accurately reflects the design requirements for implementing Clustering over the WAN for this deployment?

A

Clustering over the WAN is unsupported because CUCM requires an RTT of less than 40 ms between any two subscriber nodes in the cluster.

B

Clustering over the WAN is unsupported unless all endpoints across both data centers are migrated to an external cloud PBX like Webex Calling.

C

Clustering over the WAN is supported because the RTT is under the 80 ms limit, but QoS-protected bandwidth must be provisioned for ICCS traffic.

D

Clustering over the WAN is supported, but database replication must be disabled between data center nodes to prevent Informix lockups.

Test Your Knowledge

A collaboration architect evaluates 1:1 versus 2:1 server redundancy models for a CUCM cluster supporting 20,000 registered endpoints across two sites on Large-OVA subscribers. What is the primary operational trade-off of selecting the 2:1 redundancy model instead of the 1:1 redundancy model?

A

It reduces server count and VM cost, but one backup subscriber cannot absorb the registrations of both of its primaries if they fail together.

B

The 2:1 model requires Informix database replication to run over UDP instead of TCP, increasing synchronization latency across the WAN between sites.

C

The 2:1 model prohibits Cisco Unified SRST at remote branch offices because branch phones can list only one backup subscriber.

D

The 2:1 model increases total server count and hypervisor hardware licensing while preventing automatic failover during planned maintenance.

Test Your Knowledge

A multinational financial organization requires a migration path from on-premises CUCM to cloud collaboration. The firm requires strict feature parity with its custom XML phone services, existing analog paging systems, and third-party CTI recording integrations. Which cloud deployment model best fulfills these requirements without redesigning the core dial plan?

A

On-premises single-site CUCM deployment without external WAN connectivity.

B

Webex Dedicated Instance connected via Webex Edge Connect or partner peering.

C

Webex Calling multi-tenant architecture using cloud-native provisioning exclusively.

D

Microsoft Teams Direct Routing without an on-premises session border controller.

Sections you finish are checked off in the contents.