2.2 Queueing Theory Models (M/M/1, M/M/s) & Little's Law Applications

Key Takeaways

  • Kendall's notation A/B/c/K/m/Z provides the standard taxonomy for queueing systems, where M denotes Markovian (Poisson arrivals or exponential service times).

  • System stability strictly requires server utilization rho < 1; as utilization approaches 100%, queue length and waiting time escalate asymptotically.

  • Little's Law (L = lambda * W and Lq = lambda * Wq) is a distribution-free conservation law applicable to any stable queueing system in steady state.

  • In single-server M/M/1 systems, expected waiting time in queue is Wq = lambda / (mu * (mu - lambda)), directly coupling throughput congestion to service variability.

  • Pooling parallel servers into a shared M/M/s queue drastically reduces average waiting time and buffer requirements compared to isolated single-server queues.

Last updated: October 2026

2.2 Queueing Theory Models (M/M/1, M/M/s) & Little's Law Applications

Queueing theory is the mathematical study of waiting lines, congestion, and service delays in stochastic operational environments. In industrial and systems engineering, queueing models evaluate manufacturing assembly cells, automated guided vehicle (AGV) dispatching, telecommunication networks, warehouse receiving docks, and healthcare delivery systems to balance the cost of providing service capacity against the cost of customer waiting or work-in-process (WIP) holding.


Architecture of Queueing Systems & Kendall's Notation

Every queueing system consists of five physical components: an arrival process from an input population, a waiting line (queue buffer), a service mechanism (servers), a service discipline, and customer departure.

To standardize queueing models, British statistician David G. Kendall developed Kendall's Notation, expressed as:

A/B/c/K/m/ZA / B / c / K / m / Z

Where:

  • AA (Arrival Process): Probability distribution of interarrival times.
    • MM: Markovian / Exponential (memoryless, Poisson arrival process)
    • DD: Deterministic (constant interarrival times)
    • EkE_k: Erlang-kk distribution
    • GG: General continuous probability distribution
  • BB (Service Time Distribution): Probability distribution of service durations.
    • MM: Exponential service times
    • DD: Deterministic / constant service times
    • GG: General service time distribution
  • cc (Number of Parallel Servers): Integer count (ss or c≥1c \ge 1) of identical, independent service channels operating in parallel.
  • KK (System Capacity): Maximum permissible number of customers in the system simultaneously (in queue plus in service). When omitted, default is ∞\infty.
  • mm (Calling Population Size): Size of the customer source pool. When omitted, default is ∞\infty.
  • ZZ (Queue Discipline): The priority order for service admission. Common disciplines include:
    • FIFO (First-In, First-Out): Standard industrial queue discipline.
    • LIFO (Last-In, First-Out): Found in inventory stackers or bucket elevators.
    • SIRO (Service In Random Order): Used in communication packet routing.
    • PRI (Priority Service): Preemptive or non-preemptive urgent dispatch.

When the last three parameters are omitted (e.g., M/M/1M/M/1 or M/M/sM/M/s), the system implies an infinite queue capacity (K=∞K = \infty), an infinite calling population (m=∞m = \infty), and a FIFO discipline.


Stochastic Foundations: Poisson Processes & Exponential Distributions

Most classical queueing formulations assume Markovian (MM) dynamics governed by the memoryless property.

1. The Poisson Arrival Process

Let λ\lambda denote the mean arrival rate (customers/hour, parts/minute). The probability that exactly nn arrivals occur within an observation time window tt is given by the Poisson distribution:

P(N(t)=n)=(λt)ne−λtn!for n=0,1,2,…P(N(t) = n) = \frac{(\lambda t)^n e^{-\lambda t}}{n!} \quad \text{for } n = 0, 1, 2, \dots

  • Expected arrivals in interval tt: E[N(t)]=λtE[N(t)] = \lambda t
  • Variance of arrivals: Var(N(t))=λt\text{Var}(N(t)) = \lambda t

If the arrival counting process is Poisson, the interarrival times (TaT_a) between successive arrivals are independently and identically distributed (i.i.d.) according to an exponential distribution with parameter λ\lambda:

f(t)=λe−λt(t≥0),F(t)=1−e−λtf(t) = \lambda e^{-\lambda t} \quad (t \ge 0), \quad F(t) = 1 - e^{-\lambda t} E[Ta]=1λ,Var(Ta)=1λ2E[T_a] = \frac{1}{\lambda}, \quad \text{Var}(T_a) = \frac{1}{\lambda^2}

2. Exponential Service Times

Let μ\mu represent the mean service rate of a single server (parts processed per unit time when continuously busy). The mean service duration per customer is 1/μ1/\mu. The service duration TsT_s follows:

f(t)=μe−μt(t≥0),F(t)=1−e−μtf(t) = \mu e^{-\mu t} \quad (t \ge 0), \quad F(t) = 1 - e^{-\mu t}

3. The Memoryless (Markovian) Property

The exponential distribution is the only continuous distribution possessing the memoryless property:

P(T>t+s∣T>s)=P(T>t)P(T > t + s \mid T > s) = P(T > t)

Physical Meaning: The remaining time until an arrival or service completion does not depend on how much time has already elapsed. In an assembly station with exponential service times, if a job has already been in processing for 10 minutes, the probability that it requires an additional 5 minutes is identical to the probability that a freshly arrived job takes 5 minutes.


Universal Conservation Laws: Little's Law & Traffic Intensity

Traffic Intensity (Utilization Factor ρ\rho)

The traffic intensity measures the operational load placed on the service facility:

  • Single-Server System (c=1c = 1): ρ=λμ\rho = \frac{\lambda}{\mu}
  • Multi-Server System (c=sc = s servers): ρ=λsμ\rho = \frac{\lambda}{s\mu}

Stability Criterion (Ergodicity): For any infinite-capacity queue to reach a steady-state equilibrium, the service facility must be capable of processing work faster than the arrival rate:

ρ<1  ⟹  λ<sμ\rho < 1 \implies \lambda < s\mu

If ρ≥1\rho \ge 1, the queue length grows toward infinity over time (Lq→∞L_q \to \infty), resulting in system instability.

Little's Law

Formulated rigorously by John D. C. Little in 1961, Little's Law is a distribution-free fundamental theorem that relates the average inventory in a system to the throughput rate and the average flow time:

L=λWL = \lambda W Lq=λWqL_q = \lambda W_q

Where:

  • LL: Average number of customers/units in the entire system (queue + service).
  • LqL_q: Average number of customers/units waiting in the queue.
  • WW: Average time a customer spends in the entire system (sojourn time).
  • WqW_q: Average time a customer spends waiting in the queue.
  • λ\lambda: Average arrival rate of entering customers.

Fundamental Additive Decompositions

Total time in the system is the sum of waiting time in queue and processing time in service:

W=Wq+1μW = W_q + \frac{1}{\mu}

Multiplying throughout by λ\lambda yields the relationship between system and queue inventory:

L=Lq+λμ=Lq+ρ(for s=1)L = L_q + \frac{\lambda}{\mu} = L_q + \rho \quad (\text{for } s = 1) L=Lq+λμ=Lq+sρ(for s≥1)L = L_q + \frac{\lambda}{\mu} = L_q + s\rho \quad (\text{for } s \ge 1)

Little's Law applies universally regardless of arrival distribution, service distribution, or queue discipline, provided the system is stationary and stable.


Single-Server Queueing Model (M/M/1M/M/1)

The M/M/1M/M/1 queue represents a single server fed by a Poisson arrival process with rate λ\lambda and exponential service times with rate μ\mu.

State Probabilities

Using continuous-time Markov chain birth-death transition equations, the steady-state balance equations require: λPn=μPn+1  ⟹  Pn+1=ρPn\lambda P_n = \mu P_{n+1} \implies P_{n+1} = \rho P_n

Summing all probabilities to unity (∑n=0∞Pn=1\sum_{n=0}^\infty P_n = 1) yields the geometric distribution:

P0=1−ρ(Probability the system is completely empty / idle)P_0 = 1 - \rho \quad (\text{Probability the system is completely empty / idle}) Pn=(1−ρ)ρn(Probability exactly n customers are in the system)P_n = (1 - \rho)\rho^n \quad (\text{Probability exactly } n \text{ customers are in the system})

Probability of System Congestion (N≥kN \ge k)

The probability that the number of entities in the system is at least kk is:

P(N≥k)=∑n=k∞(1−ρ)ρn=ρkP(N \ge k) = \sum_{n=k}^\infty (1 - \rho)\rho^n = \rho^k

Performance Metric Closed-Form Equations

  1. Average number in system (LL): L=∑n=0∞nPn=ρ1−ρ=λμ−λL = \sum_{n=0}^\infty n P_n = \frac{\rho}{1 - \rho} = \frac{\lambda}{\mu - \lambda}

  2. Average number in queue (LqL_q): Lq=L−ρ=λμ−λ−λμ=λ2μ(μ−λ)L_q = L - \rho = \frac{\lambda}{\mu - \lambda} - \frac{\lambda}{\mu} = \frac{\lambda^2}{\mu(\mu - \lambda)}

  3. Average time in system (WW): W=Lλ=1μ−λW = \frac{L}{\lambda} = \frac{1}{\mu - \lambda}

  4. Average time in queue (WqW_q): Wq=Lqλ=λμ(μ−λ)W_q = \frac{L_q}{\lambda} = \frac{\lambda}{\mu(\mu - \lambda)}

The "Hockey-Stick" Congestion Curve

Because the denominator contains (μ−λ)=μ(1−ρ)(\mu - \lambda) = \mu(1 - \rho), as utilization ρ→1.0\rho \to 1.0, the waiting time curve escalates non-linearly. At ρ=0.50\rho = 0.50, L=1.0L = 1.0; at ρ=0.80\rho = 0.80, L=4.0L = 4.0; at ρ=0.95\rho = 0.95, L=19.0L = 19.0. Industrial systems designed to operate at near 100% utilization experience extreme queue inflation whenever stochastic variability is present.


Multi-Server Queueing Model (M/M/sM/M/s)

The M/M/sM/M/s model features a single waiting queue feeding ss parallel, identical servers, each processing at rate μ\mu. Total service capacity is sμs\mu.

Transition Dynamics

The aggregate service rate depends on the number of entities in the system nn: μn={nμif 0≤n<ssμif n≥s\mu_n = \begin{cases} n\mu & \text{if } 0 \le n < s \\ s\mu & \text{if } n \ge s \end{cases}

Probability of an Empty System (P0P_0)

P0=[∑n=0s−1(λ/μ)nn!+(λ/μ)ss!(1−λsμ)]−1P_0 = \left[ \sum_{n=0}^{s-1} \frac{(\lambda / \mu)^n}{n!} + \frac{(\lambda / \mu)^s}{s! \left(1 - \frac{\lambda}{s\mu}\right)} \right]^{-1}

Erlang C Formula (Probability of Waiting)

The probability that an arriving customer finds all ss servers occupied and must wait in the queue is given by the Erlang C formula (C(s,λ/μ)C(s, \lambda/\mu)):

P(wait)=(λ/μ)ss!(1−ρ)∑n=0s−1(λ/μ)nn!+(λ/μ)ss!(1−ρ)=P0⋅(λ/μ)ss!(1−ρ)P(\text{wait}) = \frac{\frac{(\lambda / \mu)^s}{s!(1 - \rho)}}{ \sum_{n=0}^{s-1} \frac{(\lambda / \mu)^n}{n!} + \frac{(\lambda / \mu)^s}{s!(1 - \rho)} } = P_0 \cdot \frac{(\lambda / \mu)^s}{s!(1 - \rho)}

Multi-Server Operational Metrics

  1. Average queue length (LqL_q): Lq=P0(λ/μ)sρs!(1−ρ)2=P(wait)⋅ρ1−ρL_q = \frac{P_0 (\lambda / \mu)^s \rho}{s!(1 - \rho)^2} = \frac{P(\text{wait}) \cdot \rho}{1 - \rho}

  2. Average time in queue (WqW_q): Wq=LqλW_q = \frac{L_q}{\lambda}

  3. Average time in system (WW): W=Wq+1μW = W_q + \frac{1}{\mu}

  4. Average number in system (LL): L=Lq+λμ=λWL = L_q + \frac{\lambda}{\mu} = \lambda W


Engineering Comparative Analysis: Single Fast Server vs. Server Pooling

Consider an industrial inspection facility handling an arrival rate of λ=16\lambda = 16 assemblies per hour. Management must decide between three design configurations:

  • Configuration A (Single Fast Server): One advanced automated cell with μ=20\mu = 20 parts/hour (M/M/1M/M/1).
  • Configuration B (Pooled Parallel Servers): Two standard inspection cells pooled with a shared queue, each operating at μ=10\mu = 10 parts/hour (M/M/2M/M/2).
  • Configuration C (Unpooled Dedicated Servers): Two separate lines, each with its own dedicated queue, splitting arrivals evenly so λ=8\lambda = 8 parts/hour and μ=10\mu = 10 parts/hour (2×M/M/12 \times M/M/1).

Quantitative Calculation Comparison

Configuration A: M/M/1M/M/1 with λ=16,μ=20\lambda = 16, \mu = 20

  • Utilization: ρ=1620=0.80\rho = \frac{16}{20} = 0.80
  • L=1620−16=4.00L = \frac{16}{20 - 16} = 4.00 assemblies
  • Lq=L−ρ=4.00−0.80=3.20L_q = L - \rho = 4.00 - 0.80 = 3.20 assemblies
  • W=120−16=0.25W = \frac{1}{20 - 16} = 0.25 hours =15.00= 15.00 minutes
  • Wq=3.2016=0.20W_q = \frac{3.20}{16} = 0.20 hours =12.00= 12.00 minutes

Configuration B: M/M/2M/M/2 with λ=16,μ=10,s=2\lambda = 16, \mu = 10, s = 2

  • Ratio λμ=1610=1.60\frac{\lambda}{\mu} = \frac{16}{10} = 1.60; Utilization ρ=1.602=0.80\rho = \frac{1.60}{2} = 0.80
  • Calculate P0P_0: P0=[1.600!+1.611!+1.622!(1−0.80)]−1=[1+1.6+2.562(0.20)]−1=[2.6+6.4]−1=19≈0.1111P_0 = \left[ \frac{1.6^0}{0!} + \frac{1.6^1}{1!} + \frac{1.6^2}{2!(1 - 0.80)} \right]^{-1} = \left[ 1 + 1.6 + \frac{2.56}{2(0.20)} \right]^{-1} = [2.6 + 6.4]^{-1} = \frac{1}{9} \approx 0.1111
  • Calculate P(wait)P(\text{wait}): P(wait)=P0⋅1.622!(1−0.80)=19⋅6.40=0.7111P(\text{wait}) = P_0 \cdot \frac{1.6^2}{2!(1 - 0.80)} = \frac{1}{9} \cdot 6.40 = 0.7111
  • Calculate LqL_q: Lq=P(wait)⋅ρ1−ρ=0.7111⋅0.800.20=2.844 assembliesL_q = \frac{P(\text{wait}) \cdot \rho}{1 - \rho} = \frac{0.7111 \cdot 0.80}{0.20} = 2.844 \text{ assemblies}
  • Calculate WqW_q: Wq=Lqλ=2.84416=0.1778 hours=10.67 minutesW_q = \frac{L_q}{\lambda} = \frac{2.844}{16} = 0.1778 \text{ hours} = 10.67 \text{ minutes}
  • Calculate WW: W=Wq+1μ=0.1778+0.1000=0.2778 hours=16.67 minutesW = W_q + \frac{1}{\mu} = 0.1778 + 0.1000 = 0.2778 \text{ hours} = 16.67 \text{ minutes}
  • Calculate LL: L=Lq+λμ=2.844+1.600=4.444 assembliesL = L_q + \frac{\lambda}{\mu} = 2.844 + 1.600 = 4.444 \text{ assemblies}

Configuration C: Unpooled 2×(M/M/1)2 \times (M/M/1) with λ=8,μ=10\lambda = 8, \mu = 10

  • Utilization: ρ=810=0.80\rho = \frac{8}{10} = 0.80
  • L=810−8=4.00L = \frac{8}{10 - 8} = 4.00 assemblies per line (Total across 2 lines =8.00= 8.00)
  • Lq=8210(10−8)=6420=3.20L_q = \frac{8^2}{10(10 - 8)} = \frac{64}{20} = 3.20 assemblies per line (Total across 2 lines =6.40= 6.40)
  • W=110−8=0.50W = \frac{1}{10 - 8} = 0.50 hours =30.00= 30.00 minutes
  • Wq=810(2)=0.40W_q = \frac{8}{10(2)} = 0.40 hours =24.00= 24.00 minutes

System Comparison Summary Table

Operational MetricConfig A: M/M/1M/M/1 (Fast)Config B: M/M/2M/M/2 (Pooled)Config C: 2×M/M/12 \times M/M/1 (Unpooled)
Arrival Rate (λ\lambda)16 parts/hr16 parts/hr2×82 \times 8 parts/hr
Service Rate per Server (μ\mu)20 parts/hr10 parts/hr10 parts/hr
Total Service Capacity (sμs\mu)20 parts/hr20 parts/hr20 parts/hr
System Utilization (ρ\rho)80.0%80.0%80.0%
Prob. of Queuing P(wait)P(\text{wait})80.0%71.1%80.0%
Mean Queue Waiting Time (WqW_q)12.00 min10.67 min24.00 min
Mean Service Time (1/μ1/\mu)3.00 min6.00 min6.00 min
Mean System Sojourn Time (WW)15.00 min16.67 min30.00 min
Mean Queue Length (LqL_q)3.20 parts2.84 parts6.40 parts (total)
Mean System WIP (LL)4.00 parts4.44 parts8.00 parts (total)

Key Industrial Insights

  1. The Pooling Principle: Moving from unpooled queues (Config C) to a shared pooled queue (Config B) with identical total capacity reduces average queue wait time (WqW_q) from 24.0 minutes to 10.67 minutes—a 55.5% reduction in waiting time with zero additional investment in machinery.
  2. Fast Single Server Trade-Off: Configuration A achieves the shortest overall system time (W=15.0W = 15.0 minutes) because its service time is half as long (3 min vs 6 min), but Configuration B provides lower queue waiting time (Wq=10.67W_q = 10.67 min) and provides redundancy (if one server fails, the facility continues operating at 50% capacity, whereas Config A experiences complete line shutdown).
Test Your Knowledge

An automated optical inspection station in a printed circuit board plant operates as an M/M/1 queue. Boards arrive according to a Poisson process at a rate of 18 boards per hour. The vision system inspects boards with exponentially distributed service times at an average rate of 24 boards per hour. On average, how long does a board spend waiting in the queue before its inspection begins?

A

7.5 minutes

B

10.0 minutes

C

12.5 minutes

D

15.0 minutes

Test Your Knowledge

A logistics warehouse operates 2 identical unloading docks as an M/M/2 queue. Freight trucks arrive at a rate of lambda = 3 trucks per hour. Each dock unloads trucks at a rate of mu = 2 trucks per hour. Given that the steady-state probability of an empty facility is P0 = 0.1429, what is the expected average number of trucks waiting in the queue (Lq)?

A

0.96 trucks

B

1.44 trucks

C

2.88 trucks

D

1.93 trucks

Sections you finish are checked off in the contents.