11.1 Real-Time Media Characteristics & Impairments
Key Takeaways
Real-time voice and video rely strictly on UDP and RTP encapsulation because TCP retransmission timers and head-of-line blocking introduce delays that destroy interactive bidirectional communication.
ITU-T G.114 establishes the global benchmark for one-way mouth-to-ear delay: 150 ms or less is required for high-quality conversational interaction, 150 ms to 400 ms is marginal, and delay exceeding 400 ms is unacceptable.
Packet Delay Variation (jitter) must remain at or below 30 ms; playout buffers absorb interarrival variation, but excessive jitter causes buffer underflow or packet discards from buffer overflow.
Toll-quality audio tolerates a maximum packet loss rate of 1%, whereas interactive video is far more sensitive because one lost packet can corrupt reference frames; Cisco's TelePresence guidance targets 0.05% or less.
The ITU-T G.107 E-model computes the transmission rating factor R (0 to 100 scale), which maps mathematically to a Mean Opinion Score (MOS) between 1.0 and 5.0, where G.711 achieves a benchmark MOS of 4.1 to 4.4 and G.729 achieves 3.8 to 4.0.
11.1 Real-Time Media Characteristics & Impairments
Real-time communications represent the most demanding category of enterprise IP network traffic. Unlike non-real-time data applications (such as web browsing, file transfers, and email) that depend on the Transmission Control Protocol (TCP) for guaranteed delivery via retransmission, real-time media cannot tolerate the retransmission latencies inherent in connection-oriented protocols. Understanding the transport characteristics of voice and video, the physiological thresholds of human conversation, and the mathematical models used to evaluate call quality is essential for collaboration network engineering.
1. Real-Time Media Network Characteristics
Voice and interactive video media streams possess distinct networking characteristics that set them apart from asynchronous data payloads:
+-------------------------------------------------------------------------+
| REAL-TIME MEDIA ENCAPSULATION STACK |
| |
| +--------------------+----------------+----------------+------------+ |
| | Layer 2 (Ethernet) | Layer 3 (IPv4) | Layer 4 (UDP) | L5-7 (RTP) | |
| | 14/18 Bytes | 20 Bytes | 8 Bytes | 12 Bytes | |
| +--------------------+----------------+----------------+------------+ |
| | Payload | |
| | (G.711 = | |
| | 160 Bytes) | |
| +------------+ |
+-------------------------------------------------------------------------+
UDP vs. TCP Transport
Real-time media utilizes the User Datagram Protocol (UDP) at Layer 4, paired with the Real-Time Transport Protocol (RTP, RFC 3550) at the application layer. TCP is fundamentally unsuited for live conversational media because:
- Retransmission Timers: When a packet is lost in transit, TCP halts buffer delivery, sends a duplicate acknowledgment, triggers a retransmission timeout (RTO), and resends the missing segment. In a real-time conversation, by the time a dropped packet is retransmitted and acknowledged (typically 100 ms to 300 ms later), the conversational moment has passed. Playing retransmitted voice audio out of sequence produces severe distortion and cognitive disruption.
- Head-of-Line Blocking: TCP enforces in-order byte stream delivery. A single dropped TCP segment stalls all subsequent arriving data in the receiver's TCP socket buffer until the missing segment is recovered.
- Flow Control & Congestion Avoidance: TCP dynamic window sizing (slow start, congestion avoidance, multiplicative decrease) throttles transmission rates in response to congestion, causing wild variations in packet delivery rates.
Consequently, real-time media discards lost packets at the receiver and relies on Packet Loss Concealment (PLC) algorithms rather than retransmissions.
Traffic Profiles: Voice vs. Video
- Voice Streams: Characterized by small, regular, constant bit rate (CBR) packets. For instance, the standard G.711 codec samples analog audio at 8,000 samples per second with 8 bits per sample (64 kbps). Encapsulated in 20 ms framing intervals, G.711 produces 50 packets per second, each containing 160 bytes of voice payload. With Layer 2 through Layer 4 headers (Ethernet 18B + IP 20B + UDP 8B + RTP 12B = 58 bytes of overhead), a single voice call generates a predictable, continuous 87.2 kbps stream.
- Interactive Video Streams: Highly bursty, variable bit rate (VBR) payloads. Video codecs (such as H.264 Advanced Video Coding and H.265 High Efficiency Video Coding) compress video into three primary frame types: Intra-coded frames (I-frames) representing full static images, Predictive frames (P-frames) representing motion vectors relative to prior frames, and Bidirectional predictive frames (B-frames) referencing both prior and subsequent frames. While P-frames generate modest packet bursts, an I-frame generates an instantaneous burst spanning dozens of maximum transmission unit (MTU, 1500-byte) packets within a single frame interval (33 ms at 30 fps), creating massive instantaneous queuing spikes.
2. Core Network Impairments & Thresholds
Real-time interactive collaboration is bounded by three fundamental network impairments: latency, jitter, and packet loss.
+-------------------------------------------------------------------------+
| KEY NETWORK IMPAIRMENT BUDGETS |
| |
| IMPAIRMENT RECOMMENDED TARGET UNACCEPTABLE CEILING |
| ==================== ================== ==================== |
| One-Way Latency <= 150 ms (G.114) > 400 ms |
| Jitter (PDV) <= 30 ms > 50 ms |
| Packet Loss (Voice) <= 1% > 2% - 3% |
| Packet Loss (Video) <= 0.05% (TelePresence) > 1% |
+-------------------------------------------------------------------------+
Latency (One-Way Delay)
One-way latency is the elapsed time required for a media packet to travel from the speaker's microphone to the listener's earpiece (mouth-to-ear delay). Total latency is the cumulative sum of multiple physical and algorithmic delay components:
- Codec / Algorithmic Delay: Time required by the Digital Signal Processor (DSP) or software codec to sample audio, execute compression algorithms, and build the packet. Standard 20 ms framing creates an inherent baseline delay of 20 ms, plus any algorithmic lookahead (typically 5 ms for G.729).
- Serialization Delay: Time required to clock bits onto the physical transmission medium. Calculated as: On high-speed Gigabit links, serialization delay is negligible (microseconds), but on constrained serial or low-speed WAN uplinks (e.g., T1 at 1.544 Mbps), serializing a 1,500-byte data packet requires approximately 7.8 ms.
- Propagation Delay: The physical travel time of electromagnetic waves across the physical medium (optical fiber, copper, or wireless). In fiber optics, signals propagate at roughly (about ). A transcontinental fiber path of 4,000 km incurs an inescapable propagation delay each way.
- Queuing Delay: Time a packet spends sitting in router and switch buffers waiting for an egress interface to become available. Unmanaged queuing under congestion is the primary cause of latency degradation.
- Playout Buffer Delay: Latency deliberately introduced at the receiving endpoint to smooth out jitter.
ITU-T Recommendation G.114 defines standard delay categories:
- 0 to 150 ms: High quality. Conversational dynamics feel entirely natural and instantaneous.
- 150 to 400 ms: Acceptable for international calling, but users begin experiencing hesitation, conversational overlap, and "talk-over" collisions where both parties speak simultaneously.
- Above 400 ms: Unacceptable for two-way interactive communication. Normal conversation breaks down into a formal, half-duplex "walkie-talkie" cadence.
Jitter (Packet Delay Variation - PDV)
Jitter is the statistical variance in packet transit times across a network. If Packet A departs at ms and arrives at ms, while Packet B departs at ms and arrives at ms, the transit delay has shifted from 30 ms to 50 ms, creating 20 ms of interarrival jitter.
To reconstruct smooth, isochronous analog waveforms, receiving endpoints implement a playout buffer (jitter buffer):
- Playout Buffer Mechanics: Arriving packets are held in a memory queue for a predetermined window before being dispatched to the audio hardware for decoding and rendering.
- Fixed Playout Buffer: Maintains a static time depth (e.g., 40 ms or 60 ms). If network delay variation remains within this window, all packets are played out seamlessly. However, if packet transit delay spikes beyond the buffer window, the packet arrives too late for its scheduled playout time and is discarded—an event called buffer underflow.
- Adaptive Playout Buffer: Dynamically expands or contracts buffer depth based on real-time statistical measurements of delay variation reported by RTCP. When network jitter surges, the buffer expands to prevent late packet discards (trading off higher total latency for reduced drop). When jitter subsides, the buffer slowly contracts to minimize mouth-to-ear delay.
- Buffer Overflow: Occurs when bursty packets arrive faster than the buffer can dequeue them, forcing the receiver to drop newly arriving packets.
- Industry Threshold: Jitter must remain . Jitter exceeding 30 ms severely strains playout buffers, causing dropped voice frames and audible robotic or metallic distortion.
Packet Loss
Packet loss occurs when intermediate network devices discard frames due to buffer exhaustion (tail drop), priority queue policing, CRC checksum errors, or link degradation.
- Random vs. Burst Loss: Single, isolated lost packets are readily concealed by modern codecs using Packet Loss Concealment (PLC), which synthesizes replacement pitch waveforms from surrounding frames. However, burst packet loss (where three or more consecutive packets are lost) defeats PLC, resulting in audible gaps, dropped syllables, and clipped words.
- Video Sensitivity: In video streams, packet loss is far more destructive. Because video codecs rely on inter-frame differential compression, dropping a single packet belonging to an I-frame corrupts the entire reference frame. This corruption propagates across all subsequent P-frames and B-frames for seconds, manifesting as severe macroblocking (tiling), image smearing, and video freezing until a new keyframe is received.
- Industry Thresholds: Toll-quality voice requires packet loss . Interactive video should stay far lower; Cisco's TelePresence guidance targets 0.05% or less.
3. Mean Opinion Score (MOS) & the ITU-T G.107 E-Model
Subjective human perception of audio clarity has historically been measured using the Mean Opinion Score (MOS), standardized by ITU-T Recommendation P.800. In formal MOS testing, a panel of human listeners rates speech samples on an absolute 1 to 5 scale:
| MOS Score | Quality Rating | Perceived User Experience |
|---|---|---|
| 5.0 | Excellent | Completely transparent; equivalent to face-to-face speech |
| 4.0 - 4.4 | Good | Toll quality; typical of high-grade PSTN wireline calling |
| 3.6 - 3.9 | Fair | Noticeable impairment, but effortless conversation |
| 3.0 - 3.5 | Poor | Substantial distortion; requires effort to comprehend |
| 1.0 - 2.9 | Bad | Severely degraded; communication impossible or unintelligible |
The ITU-T G.107 E-Model Framework
Because gathering human listening panels is impractical for network monitoring, the ITU-T developed the E-model (G.107), a computational model that calculates an objective Transmission Rating Factor () from physical network metrics:
Where the component parameters represent:
- (Basic Signal-to-Noise Ratio): Captures the intrinsic baseline noise of the transmission system and acoustic background environment (standard nominal value ).
- (Simultaneous Impairments): Degradation occurring simultaneously with the speech signal, such as incorrect overall loudness rating (OLR), side-tone issues, and quantization noise distortion.
- (Delay Impairments): Impairments caused directly by one-way mouth-to-ear delay and acoustic or electrical echo. Once one-way delay surpasses 150 ms, increases exponentially.
- (Equipment Impairment Factor): Quantifies the distortion introduced by low-bitrate compression codecs and network packet loss (). For uncompressed G.711 PCM, . For compressed G.729a, . As packet loss increases, climbs sharply depending on the codec's PLC robustness.
- (Advantage / Expectation Factor): A compensatory psychological factor that reflects user tolerance based on mobility or access convenience. For standard wireline connections, ; for cellular mobile connections, ; for satellite links in remote territories, .
Converting -Factor to MOS
The scalar -factor spans a theoretical scale from 0 to 100 (where values above 90 represent excellent quality). The ITU-T defines the mathematical conversion from to MOS as follows:
For :
For :
For :
(Note: On the standard narrowband E-model curve, the maximum mathematical MOS score achievable is 4.5, reflecting the realistic ceiling of human telephony satisfaction).
Benchmark Codec Performance
| Codec | Bitrate | Sampling Rate | Algorithmic Lookahead | Intrinsic | Ideal Network MOS |
|---|---|---|---|---|---|
| G.711u / G.711a | 64 kbps | 8 kHz (Narrowband) | 0 ms (20 ms frame) | 0 | 4.1 - 4.4 |
| G.729 / G.729a | 8 kbps | 8 kHz (Narrowband) | 5 ms (10 ms frame) | 11 | 3.8 - 4.0 |
| G.722 | 64 kbps | 16 kHz (Wideband) | 0 ms | 0 | 4.3 - 4.5 |
| Opus | 6 - 510 kbps | Up to 48 kHz (Fullband) | 5 ms (variable) | Dynamic | 4.2 - 4.6+ |
4. Real-Time Transport Control Protocol (RTCP) Metrics
RTP media transport is paired with RTCP (RFC 3550), an out-of-band control protocol that monitors media delivery statistics, synchronizes media streams, and conveys session participant identity.
Port Allocation and Bandwidth Rules
- Port Multiplexing: By convention, RTP media flows over an even-numbered UDP port (), while RTCP control packets flow over the adjacent odd-numbered UDP port (). Modern implementations also support RTP/RTCP multiplexing (RFC 5761) over a single UDP port to simplify firewall traversal.
- Bandwidth Ceiling: RTCP traffic is strictly throttled to 5% of the total session media bandwidth to prevent control messaging from crowding out real-time media. This 5% bandwidth is divided among participants, with 25% of the RTCP allocation reserved for active media senders and 75% allocated to receivers.
RTCP Packet Types & Key Metrics
- Sender Report (SR - Payload Type 200): Transmitted periodically by active media senders. Key fields include:
- NTP Timestamp (64 bits): Wall-clock absolute Greenwich Mean Time (seconds and fractional seconds). Used to correlate RTP streams originating from different clocks (e.g., synchronizing separate audio and video streams for lip-sync).
- RTP Timestamp (32 bits): Matches the media sampling clock in the corresponding RTP data packets.
- Cumulative Packet & Octet Counts: Total packets and bytes transmitted since session inception, allowing receivers to identify transmission gaps.
- Receiver Report (RR - Payload Type 201): Transmitted by non-sending media receivers to provide feedback on reception quality. Contains one or more Report Blocks encompassing:
- Fraction Lost (8 bits): The fraction of RTP packets lost since the previous SR/RR was sent, expressed as a fixed-point fraction with the binary point at the left edge (represented in 256ths).
- Cumulative Number of Packets Lost (24 bits): Signed total count of lost RTP packets since reception began.
- Extended Highest Sequence Number Received (32 bits): Sequence number tracking that accounts for sequence wrap-arounds.
- Interarrival Jitter (32 bits): An estimate of the statistical variance of the RTP packet interarrival time, calculated using the smoothed formula: Where is the sender RTP timestamp and is the receiver arrival timestamp.
- Last Sender Report Timestamp (LSR) & Delay Since Last SR (DLSR): Used by the sender to compute exact network Round-Trip Time (RTT): Where is the current time the sender receives the Receiver Report.
An enterprise collaboration network administrator notices that remote branch users report frequent 'talk-over' collisions where callers speak simultaneously, along with awkward conversational delays. Network monitoring reveals average one-way latency of 220 ms, packet loss of 0.2%, and jitter of 12 ms. According to ITU-T G.114 and real-time media engineering standards, which statement correctly explains this behavior?
The RTP sequence numbering is resetting prematurely, causing the receiving endpoint to drop out-of-order packets and inject artificial delay.
The 220 ms one-way delay exceeds the ITU-T G.114 150 ms guideline, so speakers talk over each other even though jitter and loss are low.
The packet loss rate of 0.2% exceeds the maximum tolerable threshold for G.711 voice media, preventing Packet Loss Concealment (PLC) from operating.
The jitter of 12 ms exceeds the 10 ms threshold for fixed jitter buffers, causing the playout buffer to discard valid packets as late arrivals.
In the ITU-T G.107 E-model transmission rating formula R = R0 - Is - Id - Ie + A, what does the parameter 'Ie' represent, and how does it change when an administrator transitions remote branch calls from G.711 to G.729a across an IP WAN?
Ie represents the psychological advantage factor; it increases to compensate for user expectations when operating over constrained bandwidth.
Ie represents the delay impairment factor; it increases because G.729a requires more serialization time across physical media.
Ie represents the simultaneous impairment factor; it decreases because G.729a uses smaller packet sizes that reduce quantization noise.
Ie is the equipment impairment factor; it rises from 0 for G.711 to 11 for G.729a because of compression distortion.
A network engineer captures RTCP traffic between two Cisco IP phones during a voice session. Which pair of fields in the RTCP Receiver Report (RR) and Sender Report (SR) enables the transmitting endpoint to calculate the network round-trip time (RTT)?
Last SR Timestamp (LSR) and Delay Since Last SR (DLSR) in the RR.
Interarrival Jitter in the RR and 64-bit NTP Timestamp in the SR exclusively.
Fraction Lost in the RR and Sender Octet Count in the SR.
Cumulative Number of Packets Lost in the RR and RTP Timestamp in the SR.
Sections you finish are checked off in the contents.