1.2 Capacity Planning, Sizing & Bandwidth Considerations

Key Takeaways

  • One Erlang represents 3,600 seconds of continuous circuit utilization, calculated as E = (BHCA * H) / 3600, where BHCA represents Busy Hour Call Attempts and H represents average holding time in seconds.

  • The Erlang B traffic model assumes blocked calls are cleared immediately, serving as the industry standard for sizing PSTN trunks, SIP trunks, and WAN call admission limits under a designated Grade of Service (GoS).

  • VoIP packets carry 40 bytes of IP/UDP/RTP headers (20-byte IP + 8-byte UDP + 12-byte RTP); untagged Ethernet adds 18 bytes (14-byte header plus 4-byte FCS), and an 802.1Q tag adds 4 more.

  • At standard 20 ms packetization, G.711 consumes 80.0 kbps at Layer 3 and 87.2 kbps on Ethernet L2, whereas G.729 consumes 24.0 kbps at Layer 3 and 31.2 kbps on Ethernet L2, delivering a 64.2% bandwidth reduction.

  • High Efficiency Video Coding (H.265 / HEVC) delivers a 40% to 50% bandwidth reduction over H.264 AVC at identical resolutions by employing Coding Tree Units (CTUs) up to 64x64 pixels and 35 intra-prediction modes (33 angular plus planar and DC).

Last updated: October 2026

1.2 Capacity Planning, Sizing & Bandwidth Considerations

Designing enterprise voice and video networks requires rigorous mathematical modeling of telecommunications traffic and granular analysis of packet-level encapsulation overhead. Collaboration engineers must accurately forecast concurrent session volume, determine PSTN and SIP trunk requirements, and allocate adequate Quality of Service (QoS) bandwidth across wide area network links.


1. Telephony Traffic Engineering: Erlang Models & BHCA

Telecommunications capacity planning quantifies user calling behavior during the most heavily utilized period of the business day: the Busy Hour.

Core Traffic Metrics

  • Busy Hour Call Attempts (BHCA): The total number of call setup attempts (both answered and unanswered) initiated during the peak 60-minute traffic window.
  • Holding Time (HH): The average call duration in seconds, measuring the period a trunk or audio channel remains engaged.
  • Erlang (EE): The dimensionless international standard unit of telecommunications traffic intensity. One Erlang represents continuous utilization of one voice channel for one hour (3,600 call-seconds).
  • Centum Call Seconds (CCS): Unit representing 100 call-seconds of traffic. Since one hour contains 3,600 seconds:

1 Erlang=36 CCS1 \text{ Erlang} = 36 \text{ CCS}

Traffic intensity in Erlangs is calculated using the formula:

E=BHCA×H3600E = \frac{\text{BHCA} \times H}{3600}

Example Calculation: If an office generates 1,200 BHCA with an average holding time of 180 seconds:

E=1200×1803600=2160003600=60 ErlangsE = \frac{1200 \times 180}{3600} = \frac{216000}{3600} = 60 \text{ Erlangs}

Erlang B vs. Erlang C Traffic Models

Selecting the correct Erlang distribution model depends on how the underlying telecommunications system handles blocked traffic during peak hours:

Traffic ModelUnderlying AssumptionPrimary Collaboration ApplicationOutput Variable Calculated
Erlang BBlocked Calls Cleared (Lost): Callers who encounter a busy signal hang up immediately and do not queue.Sizing PSTN T1/E1 trunks, SIP gateway trunks, and WAN call concurrency limits.Trunks required to achieve target Grade of Service (e.g., P.01 = 1% block).
Extended Erlang BBlocked Calls Retry: A fixed percentage of blocked callers immediately redial after receiving a busy signal.High-traffic PSTN ingress trunks with persistent retry behaviors.Adjusted trunk count compensating for caller retry multiplication.
Erlang CBlocked Calls Delayed (Queued): Blocked callers do not receive a busy tone; they wait in an infinite or finite FIFO queue.Sizing Cisco Unified Contact Center Enterprise (UCCE/PCCE) agent pools and IVR ports.Number of agents required to meet target Average Speed of Answer (ASA).

Grade of Service (GoS)

Grade of Service expresses the probability (PP) that a call attempt will be blocked due to insufficient trunk capacity during the busy hour. Common enterprise engineering thresholds include:

  • P.01 (1% blocking): Standard commercial enterprise trunking design. Out of 100 call attempts, no more than 1 will encounter a busy tone.
  • P.001 (0.1% blocking): Mission-critical or emergency response trunking design.

2. Voice Packet Overhead & Bandwidth Calculation

A Voice over IP (VoIP) stream consists of digitized voice samples packed into RTP packets, encapsulated across UDP, IP, and the underlying data link layer.

Protocol Encapsulation Breakdown

+-----------------------------------------------------------------------------------------+
|                                 ETHERNET VoIP PACKET                                    |
+-----------------------+-------------------+------------------+---------------+----------+
| Layer 2 Frame Header  | IPv4 Header       | UDP Header       | RTP Header    | Audio    |
| Dest MAC + Src MAC +  | 20 Bytes          | 8 Bytes          | 12 Bytes      | Payload  |
| Type + 4-byte FCS     | (40 B for IPv6)   |                  |               | (e.g.,   |
| = 18 Bytes            |                   |                  |               | 160 B)   |
+-----------------------+-------------------+------------------+---------------+----------+
|                       |<------------------ 40-BYTE L3/L4/RTP --------------->|          |
|<--------------------------- TOTAL PACKET SIZE: 218 BYTES ------------------------------>|
  1. Layer 2 Ethernet Overhead: Ethernet II framing adds a 14-byte header (6-byte destination MAC + 6-byte source MAC + 2-byte EtherType) and a 4-byte Frame Check Sequence (FCS) trailer, totaling 18 bytes of data-link overhead. Where an 802.1Q VLAN tag is present (for example, on a trunk link), it adds 4 more bytes for 22 bytes; G.711 at 20 ms then needs 88.8 kbps instead of 87.2 kbps.
  2. Layer 3 IPv4 Header: Standard IPv4 header without options consumes 20 bytes (IPv6 consumes 40 bytes).
  3. Layer 4 UDP Header: User Datagram Protocol consumes 8 bytes.
  4. Real-time Transport Protocol (RTP) Header: Standard RTP header consumes 12 bytes.

Total combined IP/UDP/RTP header overhead at Layer 3 and Layer 4:

Header Overhead=20 (IP)+8 (UDP)+12 (RTP)=40 bytes\text{Header Overhead} = 20 \text{ (IP)} + 8 \text{ (UDP)} + 12 \text{ (RTP)} = 40 \text{ bytes}

Packetization Interval and Packet Rate

VoIP endpoints sample analog signals and package them at fixed transmission intervals (packetization intervals), most commonly 20 ms (50 packets per second), 10 ms (100 packets per second), or 30 ms (33.33 packets per second):

Packets Per Second (pps)=1Packetization Interval (seconds)\text{Packets Per Second (pps)} = \frac{1}{\text{Packetization Interval (seconds)}}

Voice Payload (bytes)=Codec Bitrate (bps)×Packetization Interval (s)8 bits/byte\text{Voice Payload (bytes)} = \frac{\text{Codec Bitrate (bps)} \times \text{Packetization Interval (s)}}{8 \text{ bits/byte}}

Example (G.711 at 20 ms):

Voice Payload=64000×0.0208=160 bytes\text{Voice Payload} = \frac{64000 \times 0.020}{8} = 160 \text{ bytes}

Example (G.729 at 20 ms):

Voice Payload=8000×0.0208=20 bytes\text{Voice Payload} = \frac{8000 \times 0.020}{8} = 20 \text{ bytes}

Calculating Bitrates

  • Layer 3 Packet Size: Payload Bytes+40 bytes\text{Payload Bytes} + 40 \text{ bytes}
  • Layer 3 Bandwidth (bps): L3 Packet Size (bytes)×8 bits/byte×pps\text{L3 Packet Size (bytes)} \times 8 \text{ bits/byte} \times \text{pps}
  • Layer 2 Frame Size: L3 Packet Size+18 bytes\text{L3 Packet Size} + 18 \text{ bytes} (untagged Ethernet; add 4 bytes for 802.1Q)
  • Layer 2 Bandwidth (bps): L2 Frame Size (bytes)×8 bits/byte×pps\text{L2 Frame Size (bytes)} \times 8 \text{ bits/byte} \times \text{pps}

3. Audio Codec Bandwidth Comparison

The following table details packet dimensions and bandwidth requirements for prominent enterprise codecs across standard Ethernet infrastructure:

CodecAlgorithmic BitratePacketization IntervalPayload SizePackets Per SecondLayer 3 BandwidthEthernet L2 Bandwidth (untagged, 18 bytes)
G.711 (μ\mu-law / A-law)64.0 kbps20 ms160 bytes50 pps80.0 kbps87.2 kbps
G.711 (μ\mu-law / A-law)64.0 kbps10 ms80 bytes100 pps96.0 kbps110.4 kbps
G.711 (μ\mu-law / A-law)64.0 kbps30 ms240 bytes33.33 pps74.7 kbps79.5 kbps
G.729 / G.729a8.0 kbps20 ms20 bytes50 pps24.0 kbps31.2 kbps
G.722 (Wideband)64.0 kbps20 ms160 bytes50 pps80.0 kbps87.2 kbps
Opus (Speech Mode)~20.0 kbps (adaptive)20 ms50 bytes50 pps36.0 kbps43.2 kbps
iLBC (Mode 20)15.2 kbps20 ms38 bytes50 pps31.2 kbps38.4 kbps
iLBC (Mode 30)13.33 kbps30 ms50 bytes33.33 pps24.0 kbps28.8 kbps

Header Compression Note: Compressed RTP (cRTP, RFC 2508) compresses the 40-byte IP/UDP/RTP header down to 2 or 4 bytes over point-to-point serial WAN links (such as T1/E1 PPP/HDLC). However, cRTP is CPU-intensive and is not supported across standard Ethernet, MPLS, or IPsec encapsulated tunnels.


4. Video Bandwidth Requirements: H.264 AVC vs. H.265 HEVC

Video endpoints generate variable bitrate (VBR) or constrained bitrate (CBR) RTP streams that consume significantly greater bandwidth than voice streams.

Video Resolution and Bandwidth Matrix

Video FormatResolution (Pixels)Frame RateH.264 AVC Target BandwidthH.265 HEVC Target BandwidthTypical Use Case
720p HD1280×7201280 \times 72030 fps1.0 Mbps to 1.5 Mbps600 kbps to 900 kbpsCisco Webex Desk Series, standard desktop video
1080p Full HD1920×10801920 \times 108030 fps2.5 Mbps to 4.0 Mbps1.5 Mbps to 2.2 MbpsCisco Room Kit, Board Series, conference rooms
1080p60 High Motion1920×10801920 \times 108060 fps4.0 Mbps to 6.0 Mbps2.5 Mbps to 3.5 MbpsImmersive Telepresence, high-motion data sharing
4K Ultra HD3840×21603840 \times 216030 fps12.0 Mbps to 18.0 Mbps6.0 Mbps to 10.0 MbpsUltra-high-definition presentation and main video

Architectural Efficiency: H.264 vs. H.265

High Efficiency Video Coding (H.265 / HEVC) delivers approximately 40% to 50% bandwidth savings over H.264 Advanced Video Coding (AVC) at equivalent subjective visual quality:

  • Block Partitioning: H.264 relies on rigid 16×1616 \times 16 pixel macroblocks. H.265 replaces macroblocks with flexible Coding Tree Units (CTUs) ranging up to 64×6464 \times 64 pixels, allowing efficient compression of expansive background regions.
  • Intra-Prediction Precision: H.264 uses up to 9 intra-prediction modes (for 4x4 luma blocks); H.265 defines 35 (33 angular plus planar and DC), which predicts detail more accurately and leaves less residual to encode.
  • Network Impact: Upgrading video endpoints to H.265 enables enterprises to route 1080p30 video over WAN links provisioned for 1.8 Mbps rather than 3.5 Mbps.

5. Step-by-Step Worked Bandwidth Calculation

Problem Statement

An enterprise branch office connects to corporate headquarters across an MPLS WAN. The branch requires concurrent capacity for 100 active voice calls.

  • Calculate the required Layer 2 Ethernet WAN bandwidth for 100 concurrent G.711 calls at 20 ms packetization.
  • Calculate the required Layer 2 Ethernet WAN bandwidth for 100 concurrent G.729 calls at 20 ms packetization.
  • Factor in a 5% signaling buffer (SIP and RTCP) and a 20% QoS Low Latency Queuing (LLQ) headroom reserve to determine the final provisioned Priority Queue size.

Step-by-Step Solution

Part A: Sizing for 100 G.711 Calls (20 ms)

  1. Determine Payload: G.711 produces 64,000 bps. Over 20 ms, payload is: 64000×0.0208=160 bytes\frac{64000 \times 0.020}{8} = 160 \text{ bytes}
  2. Calculate L3 Packet Size: Add 40 bytes for IP/UDP/RTP headers: 160+40=200 bytes160 + 40 = 200 \text{ bytes}
  3. Calculate L2 Frame Size: Add 18 bytes of untagged Ethernet overhead: 200+18=218 bytes200 + 18 = 218 \text{ bytes}
  4. Calculate Single Call L2 Bitrate: At 50 packets per second: 218 bytes×8 bits/byte×50 pps=87,200 bps=87.2 kbps218 \text{ bytes} \times 8 \text{ bits/byte} \times 50 \text{ pps} = 87,200 \text{ bps} = 87.2 \text{ kbps}
  5. Aggregate Media Bandwidth (100 Calls): 100×87.2 kbps=8,720 kbps=8.72 Mbps100 \times 87.2 \text{ kbps} = 8,720 \text{ kbps} = 8.72 \text{ Mbps}
  6. Add 5% Signaling (SIP/RTCP): 8.72 Mbps×1.05=9.156 Mbps8.72 \text{ Mbps} \times 1.05 = 9.156 \text{ Mbps}
  7. Apply 20% QoS Headroom Safety Margin: 9.156 Mbps×1.20=10.987 Mbps≈11.0 Mbps9.156 \text{ Mbps} \times 1.20 = 10.987 \text{ Mbps} \approx 11.0 \text{ Mbps}

Part B: Sizing for 100 G.729 Calls (20 ms)

  1. Determine Payload: G.729 produces 8,000 bps. Over 20 ms, payload is: 8000×0.0208=20 bytes\frac{8000 \times 0.020}{8} = 20 \text{ bytes}
  2. Calculate L3 Packet Size: Add 40 bytes for IP/UDP/RTP headers: 20+40=60 bytes20 + 40 = 60 \text{ bytes}
  3. Calculate L2 Frame Size: Add 18 bytes of untagged Ethernet overhead: 60+18=78 bytes60 + 18 = 78 \text{ bytes}
  4. Calculate Single Call L2 Bitrate: At 50 packets per second: 78 bytes×8 bits/byte×50 pps=31,200 bps=31.2 kbps78 \text{ bytes} \times 8 \text{ bits/byte} \times 50 \text{ pps} = 31,200 \text{ bps} = 31.2 \text{ kbps}
  5. Aggregate Media Bandwidth (100 Calls): 100×31.2 kbps=3,120 kbps=3.12 Mbps100 \times 31.2 \text{ kbps} = 3,120 \text{ kbps} = 3.12 \text{ Mbps}
  6. Add 5% Signaling (SIP/RTCP): 3.12 Mbps×1.05=3.276 Mbps3.12 \text{ Mbps} \times 1.05 = 3.276 \text{ Mbps}
  7. Apply 20% QoS Headroom Safety Margin: 3.276 Mbps×1.20=3.931 Mbps≈3.93 Mbps3.276 \text{ Mbps} \times 1.20 = 3.931 \text{ Mbps} \approx 3.93 \text{ Mbps}

Summary of Results

Transitioning 100 concurrent calls from G.711 to G.729 reduces the required priority queue allocation from 11.0 Mbps down to 3.93 Mbps, generating an immediate WAN bandwidth conservation of 64.2%.

Loading diagram...
VoIP Protocol Encapsulation and Bandwidth Amplification
Test Your Knowledge

A network engineer is sizing a SIP trunk to handle an estimated 1,800 Busy Hour Call Attempts (BHCA) with an average holding time of 180 seconds per call. If the target Grade of Service (GoS) is P.01 using the Erlang B traffic model, how many Erlangs of traffic must the trunk support during the busy hour?

A

60 Erlangs

B

30 Erlangs

C

10 Erlangs

D

90 Erlangs

Test Your Knowledge

When calculating WAN bandwidth for a G.711 audio stream using standard 20 ms packetization over untagged Ethernet (18 bytes of Layer 2 overhead), what is the resulting single-call Layer 3 bandwidth and single-call Layer 2 Ethernet bandwidth?

A

96.0 kbps Layer 3 bandwidth and 110.4 kbps Layer 2 bandwidth.

B

80.0 kbps Layer 3 bandwidth and 87.2 kbps Layer 2 bandwidth.

C

64.0 kbps Layer 3 bandwidth and 80.0 kbps Layer 2 bandwidth.

D

24.0 kbps Layer 3 bandwidth and 31.2 kbps Layer 2 bandwidth.

Test Your Knowledge

An enterprise video collaboration deployment upgrades its video endpoints from H.264 AVC to H.265 HEVC. What architectural mechanism in H.265 enables up to a 50% reduction in required bandwidth while maintaining equivalent 1080p video quality?

A

Converting all video color spaces from 4:2:0 chroma subsampling to uncompressed 4:4:4 color arrays.

B

Coding Tree Units of up to 64x64 pixels with 35 intra-prediction modes, replacing fixed 16x16 macroblocks.

C

Replacing UDP transport with TCP windowing so that lost video packets are retransmitted rather than concealed.

D

Lowering the video frame rate from 30 fps to 15 fps during high-motion scenes.

Sections you finish are checked off in the contents.