2.2 Queueing Theory Models (M/M/1, M/M/s) & Little's Law Applications
Key Takeaways
Kendall's notation A/B/c/K/m/Z provides the standard taxonomy for queueing systems, where M denotes Markovian (Poisson arrivals or exponential service times).
System stability strictly requires server utilization rho < 1; as utilization approaches 100%, queue length and waiting time escalate asymptotically.
Little's Law (L = lambda * W and Lq = lambda * Wq) is a distribution-free conservation law applicable to any stable queueing system in steady state.
In single-server M/M/1 systems, expected waiting time in queue is Wq = lambda / (mu * (mu - lambda)), directly coupling throughput congestion to service variability.
Pooling parallel servers into a shared M/M/s queue drastically reduces average waiting time and buffer requirements compared to isolated single-server queues.
2.2 Queueing Theory Models (M/M/1, M/M/s) & Little's Law Applications
Queueing theory is the mathematical study of waiting lines, congestion, and service delays in stochastic operational environments. In industrial and systems engineering, queueing models evaluate manufacturing assembly cells, automated guided vehicle (AGV) dispatching, telecommunication networks, warehouse receiving docks, and healthcare delivery systems to balance the cost of providing service capacity against the cost of customer waiting or work-in-process (WIP) holding.
Architecture of Queueing Systems & Kendall's Notation
Every queueing system consists of five physical components: an arrival process from an input population, a waiting line (queue buffer), a service mechanism (servers), a service discipline, and customer departure.
To standardize queueing models, British statistician David G. Kendall developed Kendall's Notation, expressed as:
Where:
- (Arrival Process): Probability distribution of interarrival times.
- : Markovian / Exponential (memoryless, Poisson arrival process)
- : Deterministic (constant interarrival times)
- : Erlang- distribution
- : General continuous probability distribution
- (Service Time Distribution): Probability distribution of service durations.
- : Exponential service times
- : Deterministic / constant service times
- : General service time distribution
- (Number of Parallel Servers): Integer count ( or ) of identical, independent service channels operating in parallel.
- (System Capacity): Maximum permissible number of customers in the system simultaneously (in queue plus in service). When omitted, default is .
- (Calling Population Size): Size of the customer source pool. When omitted, default is .
- (Queue Discipline): The priority order for service admission. Common disciplines include:
- FIFO (First-In, First-Out): Standard industrial queue discipline.
- LIFO (Last-In, First-Out): Found in inventory stackers or bucket elevators.
- SIRO (Service In Random Order): Used in communication packet routing.
- PRI (Priority Service): Preemptive or non-preemptive urgent dispatch.
When the last three parameters are omitted (e.g., or ), the system implies an infinite queue capacity (), an infinite calling population (), and a FIFO discipline.
Stochastic Foundations: Poisson Processes & Exponential Distributions
Most classical queueing formulations assume Markovian () dynamics governed by the memoryless property.
1. The Poisson Arrival Process
Let denote the mean arrival rate (customers/hour, parts/minute). The probability that exactly arrivals occur within an observation time window is given by the Poisson distribution:
- Expected arrivals in interval :
- Variance of arrivals:
If the arrival counting process is Poisson, the interarrival times () between successive arrivals are independently and identically distributed (i.i.d.) according to an exponential distribution with parameter :
2. Exponential Service Times
Let represent the mean service rate of a single server (parts processed per unit time when continuously busy). The mean service duration per customer is . The service duration follows:
3. The Memoryless (Markovian) Property
The exponential distribution is the only continuous distribution possessing the memoryless property:
Physical Meaning: The remaining time until an arrival or service completion does not depend on how much time has already elapsed. In an assembly station with exponential service times, if a job has already been in processing for 10 minutes, the probability that it requires an additional 5 minutes is identical to the probability that a freshly arrived job takes 5 minutes.
Universal Conservation Laws: Little's Law & Traffic Intensity
Traffic Intensity (Utilization Factor )
The traffic intensity measures the operational load placed on the service facility:
- Single-Server System ():
- Multi-Server System ( servers):
Stability Criterion (Ergodicity): For any infinite-capacity queue to reach a steady-state equilibrium, the service facility must be capable of processing work faster than the arrival rate:
If , the queue length grows toward infinity over time (), resulting in system instability.
Little's Law
Formulated rigorously by John D. C. Little in 1961, Little's Law is a distribution-free fundamental theorem that relates the average inventory in a system to the throughput rate and the average flow time:
Where:
- : Average number of customers/units in the entire system (queue + service).
- : Average number of customers/units waiting in the queue.
- : Average time a customer spends in the entire system (sojourn time).
- : Average time a customer spends waiting in the queue.
- : Average arrival rate of entering customers.
Fundamental Additive Decompositions
Total time in the system is the sum of waiting time in queue and processing time in service:
Multiplying throughout by yields the relationship between system and queue inventory:
Little's Law applies universally regardless of arrival distribution, service distribution, or queue discipline, provided the system is stationary and stable.
Single-Server Queueing Model ()
The queue represents a single server fed by a Poisson arrival process with rate and exponential service times with rate .
State Probabilities
Using continuous-time Markov chain birth-death transition equations, the steady-state balance equations require:
Summing all probabilities to unity () yields the geometric distribution:
Probability of System Congestion ()
The probability that the number of entities in the system is at least is:
Performance Metric Closed-Form Equations
-
Average number in system ():
-
Average number in queue ():
-
Average time in system ():
-
Average time in queue ():
The "Hockey-Stick" Congestion Curve
Because the denominator contains , as utilization , the waiting time curve escalates non-linearly. At , ; at , ; at , . Industrial systems designed to operate at near 100% utilization experience extreme queue inflation whenever stochastic variability is present.
Multi-Server Queueing Model ()
The model features a single waiting queue feeding parallel, identical servers, each processing at rate . Total service capacity is .
Transition Dynamics
The aggregate service rate depends on the number of entities in the system :
Probability of an Empty System ()
Erlang C Formula (Probability of Waiting)
The probability that an arriving customer finds all servers occupied and must wait in the queue is given by the Erlang C formula ():
Multi-Server Operational Metrics
-
Average queue length ():
-
Average time in queue ():
-
Average time in system ():
-
Average number in system ():
Engineering Comparative Analysis: Single Fast Server vs. Server Pooling
Consider an industrial inspection facility handling an arrival rate of assemblies per hour. Management must decide between three design configurations:
- Configuration A (Single Fast Server): One advanced automated cell with parts/hour ().
- Configuration B (Pooled Parallel Servers): Two standard inspection cells pooled with a shared queue, each operating at parts/hour ().
- Configuration C (Unpooled Dedicated Servers): Two separate lines, each with its own dedicated queue, splitting arrivals evenly so parts/hour and parts/hour ().
Quantitative Calculation Comparison
Configuration A: with
- Utilization:
- assemblies
- assemblies
- hours minutes
- hours minutes
Configuration B: with
- Ratio ; Utilization
- Calculate :
- Calculate :
- Calculate :
- Calculate :
- Calculate :
- Calculate :
Configuration C: Unpooled with
- Utilization:
- assemblies per line (Total across 2 lines )
- assemblies per line (Total across 2 lines )
- hours minutes
- hours minutes
System Comparison Summary Table
| Operational Metric | Config A: (Fast) | Config B: (Pooled) | Config C: (Unpooled) |
|---|---|---|---|
| Arrival Rate () | 16 parts/hr | 16 parts/hr | parts/hr |
| Service Rate per Server () | 20 parts/hr | 10 parts/hr | 10 parts/hr |
| Total Service Capacity () | 20 parts/hr | 20 parts/hr | 20 parts/hr |
| System Utilization () | 80.0% | 80.0% | 80.0% |
| Prob. of Queuing | 80.0% | 71.1% | 80.0% |
| Mean Queue Waiting Time () | 12.00 min | 10.67 min | 24.00 min |
| Mean Service Time () | 3.00 min | 6.00 min | 6.00 min |
| Mean System Sojourn Time () | 15.00 min | 16.67 min | 30.00 min |
| Mean Queue Length () | 3.20 parts | 2.84 parts | 6.40 parts (total) |
| Mean System WIP () | 4.00 parts | 4.44 parts | 8.00 parts (total) |
Key Industrial Insights
- The Pooling Principle: Moving from unpooled queues (Config C) to a shared pooled queue (Config B) with identical total capacity reduces average queue wait time () from 24.0 minutes to 10.67 minutes—a 55.5% reduction in waiting time with zero additional investment in machinery.
- Fast Single Server Trade-Off: Configuration A achieves the shortest overall system time ( minutes) because its service time is half as long (3 min vs 6 min), but Configuration B provides lower queue waiting time ( min) and provides redundancy (if one server fails, the facility continues operating at 50% capacity, whereas Config A experiences complete line shutdown).
An automated optical inspection station in a printed circuit board plant operates as an M/M/1 queue. Boards arrive according to a Poisson process at a rate of 18 boards per hour. The vision system inspects boards with exponentially distributed service times at an average rate of 24 boards per hour. On average, how long does a board spend waiting in the queue before its inspection begins?
7.5 minutes
10.0 minutes
12.5 minutes
15.0 minutes
A logistics warehouse operates 2 identical unloading docks as an M/M/2 queue. Freight trucks arrive at a rate of lambda = 3 trucks per hour. Each dock unloads trucks at a rate of mu = 2 trucks per hour. Given that the steady-state probability of an empty facility is P0 = 0.1429, what is the expected average number of trucks waiting in the queue (Lq)?
0.96 trucks
1.44 trucks
2.88 trucks
1.93 trucks
Sections you finish are checked off in the contents.