10.4 Performance Measurement: Productivity, KPIs, Wage Systems, and the Balanced Scorecard
Key Takeaways
- Productivity is always output divided by input: partial-factor productivity divides by one resource (labor hours, machine hours, kilowatt-hours), while multifactor and total-factor productivity divide a common-currency output by the summed value of several or all inputs.
- Utilization, efficiency, and effectiveness measure different things: utilization is the fraction of available time a resource ran, efficiency compares actual output to standard output, and Overall Equipment Effectiveness multiplies Availability, Performance, and Quality into a single ratio.
- A usable KPI set balances lagging outcome measures against leading process measures across the Safety, Quality, Delivery, Cost, and People families, and each KPI must name an owner, a formula, a data source, a frequency, and a target.
- Wage payment plans differ in how the productivity gain is split: straight piece rate and the standard hour plan give the worker 100% of the gain, Halsey 50-50 splits it evenly, and the Rowan plan pays a bonus fraction that shrinks as efficiency rises.
- Kaplan and Norton's Balanced Scorecard translates strategy into four linked perspectives — Financial, Customer, Internal Business Process, and Learning and Growth — where learning drives process, process drives customer outcomes, and customer outcomes drive financial results.
10.4 Performance Measurement: Productivity, KPIs, Wage Systems, and the Balanced Scorecard
The NCEES FE Industrial and Systems specification lists performance measurement — key performance indicators, productivity, wage scales, balanced scorecard, and customer satisfaction — as a distinct sub-topic of Engineering Management. It is the discipline that converts operational activity into numbers an executive can act on. Industrial engineers own these measures because a badly chosen metric silently rewards the wrong behavior: reward machine utilization and you will get overproduction; reward pieces per hour with no quality gate and you will get scrap.
1. Productivity Measurement
Productivity is a ratio of output produced to input consumed over a defined period:
The measure is only meaningful when the period, the boundary, and the units are stated. Three families are distinguished by how many inputs appear in the denominator.
Partial-Factor Productivity
A single resource sits in the denominator, so output and input may be in physical units:
Partial measures are easy to collect and easy to game. Labor productivity rises when a plant automates, even if the total cost per unit rises, because the substituted capital never appears in the denominator.
Multifactor and Total-Factor Productivity
When several inputs are combined, they must be converted to a common currency:
Total-factor productivity (TFP) extends the denominator to every purchased input. A dimensionless MFP or TFP above 1.0 means the operation generates more output value than it consumes in input value.
Productivity Index and Growth
Comparisons across periods use an index against a base period:
Exam Watchout — Deflate Before Comparing: If output is measured in dollars, a 5% selling-price increase inflates measured productivity even though nothing physical improved. Multi-period productivity comparisons must use constant (real) dollars for both output and inputs, applying the same base-year price deflator used in the inflation analysis of engineering economics.
2. Utilization, Efficiency, Effectiveness, and OEE
These three words are not interchangeable, and the FE exam tests the distinction directly.
| Metric | Formula | Question It Answers |
|---|---|---|
| Utilization | $\dfrac{\text{Time the resource actually ran}}{\text{Time the resource was available}}$ | Was the asset busy? |
| Efficiency | $\dfrac{\text{Actual output}}{\text{Standard (expected) output}}$ | Did it run at the standard rate? |
| Effectiveness | $\dfrac{\text{Output that met the objective}}{\text{Output planned}}$ | Did it produce the right result? |
A press can run 95% of the shift (high utilization) at 60% of standard speed (poor efficiency) producing parts that fail final gauge (zero effectiveness).
Overall Equipment Effectiveness (OEE)
OEE is the standard composite equipment metric in lean and TPM programs. It multiplies three independent loss categories:
- Availability ($A$) captures downtime losses: breakdowns, changeovers, starvation. Planned production time already excludes scheduled non-production such as breaks and planned maintenance.
- Performance ($P$) captures speed losses: minor stops, idling, running below rated rate. $P$ can never legitimately exceed 1.0; a value above 1.0 means the "ideal" cycle time is wrong.
- Quality ($Q$) captures defect losses: scrap and rework, including startup rejects.
A useful algebraic shortcut collapses all three:
World-class discrete manufacturing OEE is commonly benchmarked near 85% ($A \approx 90%$, $P \approx 95%$, $Q \approx 99%$); typical plants land between 45% and 65%.
Two related composites appear in maintenance-heavy plants: TEEP (Total Effective Equipment Performance) multiplies OEE by the schedule-loading fraction to measure against all 168 calendar hours, while OOE uses total operating time rather than planned production time in the availability term.
3. Designing Key Performance Indicators (KPIs)
A key performance indicator is a quantified measure tied to a specific objective and used to trigger action. A metric that no one will act on is reporting overhead, not a KPI.
The Anatomy of a Defined KPI
Every KPI on a plant scorecard must specify seven attributes:
- Name and objective it supports
- Formula, written explicitly with numerator and denominator
- Unit of measure
- Data source and collection method (MES tag, ERP transaction, manual count)
- Reporting frequency (per shift, daily, weekly, monthly)
- Owner — a single named accountable person
- Target and threshold for action
Leading vs. Lagging Indicators
- Lagging indicators report a completed outcome: recordable injury rate, scrap dollars, on-time delivery percentage, gross margin. They are accurate but arrive too late to change the result.
- Leading indicators predict the outcome and can still be influenced: near-miss reports filed, preventive maintenance compliance, first-pass yield at the upstream operation, training hours completed, supplier PPM at receiving.
A balanced operations review pairs each lagging outcome with the leading measures believed to drive it.
The SQDCP Families
Most manufacturing scorecards organize KPIs into five families, deliberately kept in this order so safety is discussed first:
| Family | Representative KPIs |
|---|---|
| Safety | TRIR, DART rate, near-miss reports per 100 employees, ergonomic assessments closed |
| Quality | First-pass yield, defects per million opportunities, cost of poor quality as % of sales, supplier PPM, warranty claim rate |
| Delivery | On-time in-full (OTIF), schedule attainment, manufacturing lead time, inventory turns, fill rate |
| Cost | Unit conversion cost, labor cost per unit, OEE, scrap and rework dollars, energy cost per unit |
| People | Absenteeism, voluntary turnover, cross-training matrix coverage, suggestions implemented per employee |
Inventory turns, a recurring exam quantity, is defined as:
Common KPI Pathologies
- Vanity metrics: Impressive numbers with no decision attached (total pieces produced since 1998).
- Metric conflict: Rewarding departmental machine utilization while also demanding low WIP; the two cannot both be maximized.
- Denominator drift: Changing the population or period mid-year so the trend line becomes meaningless.
- Goodhart's law: "When a measure becomes a target, it ceases to be a good measure." Measuring only pieces per hour reliably produces pieces per hour — and scrap. Pair every rate metric with a quality gate.
4. Wage Scales, Job Evaluation, and Incentive Plans
Performance measurement and compensation are directly coupled: the wage plan determines whose behavior the measurement actually changes.
Establishing the Wage Scale: Job Evaluation
A wage scale is the structure of pay grades assigned to jobs. Four classical job-evaluation methods set it:
| Method | Basis | Mechanics |
|---|---|---|
| Ranking | Whole job, qualitative | Jobs ordered simplest to most demanding. Fast; no scale of difference. |
| Classification (Grading) | Whole job, qualitative | Pre-written grade descriptions; each job slotted into the best-fitting grade (the U.S. federal GS system). |
| Point Factor | Job parts, quantitative | Compensable factors (skill, effort, responsibility, working conditions) are weighted and scored; total points map to a pay grade. Most defensible and most common. |
| Factor Comparison | Job parts, quantitative | Benchmark jobs are ranked factor by factor and their existing wage is allocated across factors, producing a monetary scale. |
Point totals are plotted against market wage data to fit a wage curve; jobs above the line are red-circled (pay frozen until the grade catches up) and jobs below are green-circled (scheduled for increase).
Wage Payment Plans
Let $R$ = hourly base rate, $H$ = hours actually worked, $S$ = standard hours earned (pieces produced multiplied by standard time per piece), and time saved $= S - H$.
| Plan | Earnings Formula | Share of the Gain |
|---|---|---|
| Day Work (hourly) | $E = H \times R$ | 0% to the worker |
| Measured Daywork | $E = H \times R$, with pay grade reset periodically from measured performance | Delayed |
| Straight Piece Rate | $E = N_{\text{pieces}} \times r_{\text{piece}}$, where $r_{\text{piece}} = R \times$ standard time per piece | 100% to the worker |
| Standard Hour Plan | $E = S \times R$ (with $H \times R$ guaranteed as a floor) | 100% to the worker |
| Halsey 50-50 | $E = H R + 0.5 (S - H) R$ | 50% to the worker |
| Rowan | $E = H R + \left(\dfrac{S - H}{S}\right) H R$ | A share that decreases as efficiency rises |
| Taylor Differential Piece Rate | Low piece rate below standard, high piece rate at or above standard | Step change at standard |
| Gainsharing (Scanlon, Improshare) | Group bonus from plant-level labor-cost savings | Shared by the whole plant |
Two structural notes the exam probes:
- The guaranteed base rate. Modern piece-rate and standard-hour plans guarantee $H \times R$ regardless of output, so the incentive only ever adds to pay. The Taylor differential plan, which penalizes below-standard output, is largely of historical interest for exactly this reason.
- Rowan's self-limiting bonus. Because the bonus fraction is $(S - H)/S$, an operator who halves the standard time earns at most a 100% bonus and the marginal reward shrinks as efficiency climbs. Halsey's fixed 50% share does not self-limit, which is why loose standards are far more expensive under Halsey than under Rowan.
5. Customer Satisfaction Measurement
Internal efficiency metrics are meaningless if the customer is unhappy, so the specification pairs productivity and wage measures with customer satisfaction.
Perceptual Measures
- CSAT (Customer Satisfaction Score): Percentage of respondents selecting the top boxes on a satisfaction scale. Report the top-two-box percentage rather than a mean, because the underlying scale is ordinal.
- Net Promoter Score (NPS): From the 0-10 "how likely are you to recommend" question, classify 9-10 as Promoters, 7-8 as Passives, and 0-6 as Detractors: NPS ranges from $-100$ to $+100$ and is reported as an integer, not a percentage. Passives are counted in the denominator but never in the numerator.
- Customer Effort Score (CES): How much work the customer had to do to get resolution; a strong predictor of repurchase in service settings.
- SERVQUAL gap model: Measures the difference between customer expectation and perception across reliability, assurance, tangibles, empathy, and responsiveness.
Objective (Behavioral) Measures
Perceptual scores lag; objective service metrics can be measured every day:
- On-Time In-Full (OTIF): Fraction of orders delivered complete and on the promised date.
- Perfect Order Index: The product of the on-time rate, complete rate, damage-free rate, and correct-documentation rate. Because it multiplies, four separate 95% performances yield only $0.95^4 = 81.5%$ perfect orders — the same multiplicative erosion seen in series reliability.
- Warranty claim rate, return rate, and complaint rate per million units shipped.
- First-contact resolution rate and mean time to resolve.
6. The Balanced Scorecard
Robert Kaplan and David Norton introduced the Balanced Scorecard in 1992 to correct the dominance of purely financial reporting, which measures only past results and encourages short-term decisions. The scorecard translates strategy into objectives, measures, targets, and initiatives across four linked perspectives.
The Balanced Scorecard Causal Chain
┌──────────────────────────┐
│ LEARNING AND GROWTH │ Skills, information systems, culture
│ "Can we keep improving?" │ KPIs: cross-training coverage, training
└────────────┬─────────────┘ hours, employee engagement, turnover
│ enables
▼
┌──────────────────────────┐
│ INTERNAL BUSINESS PROCESS│ The processes we must excel at
│ "What must we excel at?" │ KPIs: OEE, first-pass yield, cycle time,
└────────────┬─────────────┘ schedule attainment, cost of poor quality
│ drives
▼
┌──────────────────────────┐
│ CUSTOMER │ How customers see us
│ "How do customers see us?"│ KPIs: OTIF, NPS, warranty rate, share
└────────────┬─────────────┘
│ produces
▼
┌──────────────────────────┐
│ FINANCIAL │ How we look to shareholders
│ "How do we look to owners?"│ KPIs: ROI, EVA, revenue growth, margin
└──────────────────────────┘
Key properties tested on the exam:
- The four perspectives are Financial, Customer, Internal Business Process, and Learning and Growth. "Competitor" and "Regulatory" are not scorecard perspectives.
- The perspectives are causally linked bottom-up, and this chain is drawn explicitly on a strategy map. Learning and Growth is the foundation; Financial is the outcome.
- The scorecard deliberately balances lagging financial outcomes against leading operational and human-capital drivers, and balances external (financial, customer) against internal (process, learning) views.
- Each objective carries a measure, a target, and a funded initiative; a scorecard without initiatives is a report, not a management system.
Benchmarking
Targets are set credibly by benchmarking against a reference: internal (best shift or sister plant), competitive (direct rivals), functional (best-in-class at the same function in another industry — the classic case being a hospital studying an aircraft pit crew for patient handoffs), or generic (universal processes such as order entry).
7. Step-by-Step Worked Engineering Calculations
Worked Example 10.4.1: Multifactor Productivity and Productivity Growth
Problem: A machining plant reports the following weekly figures. Output is valued at $50.00 per unit and all dollar figures are already expressed in constant base-year dollars.
| Period | Units | Labor | Material | Energy | Capital and Overhead |
|---|---|---|---|---|---|
| Baseline | 12,000 | 1,600 hr at $28.00/hr | $310,000 | $18,000 | $52,000 |
| After Kaizen | 13,200 | 1,600 hr at $28.00/hr | $335,000 | $18,500 | $52,000 |
- Compute baseline labor productivity and baseline multifactor productivity.
- Compute both measures after the Kaizen event.
- Compute the percentage growth in each, and explain the divergence.
Solution:
Step 1: Baseline
- Labor cost: $1{,}600 \times $28.00 = $44{,}800$
- Output value: $12{,}000 \times $50.00 = $600{,}000$
- Total input value: $$44{,}800 + $310{,}000 + $18{,}000 + $52{,}000 = $424{,}800$
Step 2: After Kaizen
- Output value: $13{,}200 \times $50.00 = $660{,}000$
- Total input value: $$44{,}800 + $335{,}000 + $18{,}500 + $52{,}000 = $450{,}300$
Step 3: Growth
Engineering interpretation: Labor productivity jumped 10% because the same 1,600 hours produced 1,200 more units. Multifactor productivity rose only 3.8% because material consumption grew nearly in proportion to output ($+8.1%$), so most of the extra output was bought, not earned. Reporting only the labor number would overstate the improvement by a factor of roughly 2.7. This is exactly why capital-substitution projects must be judged on MFP.
Worked Example 10.4.2: Overall Equipment Effectiveness
Problem: An injection molding press is scheduled for a 480-minute shift. Contractual breaks and a planned mold-preventive-maintenance window total 45 minutes and are excluded from planned production time. During the shift the press suffered 55 minutes of unplanned downtime (a hydraulic fault and a material starvation event). It produced 620 total shots at an ideal cycle time of 0.50 minutes per shot; 595 shots passed final inspection.
Compute Availability, Performance, Quality, and OEE, and identify the dominant loss.
Solution:
Step 1: Availability
Step 2: Performance
Step 3: Quality
Step 4: OEE Verification using the shortcut:
Step 5: Loss diagnosis: The largest single loss is Performance at 81.6% — the press ran 380 minutes but delivered only 310 minutes of ideal-cycle output, meaning 70 minutes were consumed by minor stops and slow cycles. Availability losses cost 55 minutes and quality losses cost 25 shots. Because minor stops are invisible on a downtime log, the corrective action is short-interval cycle logging at the press, not another maintenance program.
Worked Example 10.4.3: Comparing Wage Incentive Plans
Problem: The standard time for a bracket weldment is 12.0 minutes per piece, so the standard output is 5 pieces per hour. The guaranteed base rate is $R = $24.00$ per hour. In an 8-hour shift ($H = 8$), an operator completes 46 acceptable pieces; the shift standard is 40 pieces.
Compute the operator's shift earnings and the employer's direct labor cost per piece under (a) day work, (b) straight piece rate, (c) the standard hour plan, (d) Halsey 50-50, and (e) the Rowan plan.
Solution:
Step 0: Establish standard hours earned
(a) Day Work
(b) Straight Piece Rate Piece price $= R \times$ standard time per piece $= $24.00 \times 0.200 = $4.80$ per piece.
(c) Standard Hour Plan Identical to the straight piece rate — the two plans are algebraically the same whenever the piece price is derived from the base rate and the standard time.
(d) Halsey 50-50
(e) Rowan
Comparison and interpretation:
| Plan | Shift Earnings | Bonus over Day Work | Labor Cost per Piece |
|---|---|---|---|
| Day work | $192.00 | — | $4.174 |
| Halsey 50-50 | $206.40 | $14.40 | $4.487 |
| Rowan | $217.04 | $25.04 | $4.718 |
| Straight piece rate / standard hour | $220.80 | $28.80 | $4.800 |
At 115% efficiency the plans rank Halsey < Rowan < piece rate in operator earnings, and the ranking of employer unit labor cost is the mirror image. Day work shows the lowest unit labor cost here only because the operator produced 15% above standard without additional pay — a condition that does not persist, which is the entire rationale for incentive plans. Note also that at very high efficiency the Rowan bonus fraction $(S-H)/S$ approaches 1.0 and flattens, capping the employer's exposure to a loose standard, whereas Halsey's fixed 50% share and the straight piece rate do not self-limit.
A fabrication plant produces output valued at $480,000 in a week. Input costs for that week are labor $95,000, materials $210,000, energy $15,000, and capital and overhead $60,000. What is the plant's multifactor productivity?
A CNC machining center has 420 minutes of planned production time in a shift. It experiences 60 minutes of unplanned downtime. During the remaining operating time it produces 260 parts at an ideal cycle time of 1.2 minutes per part, of which 250 parts pass inspection. What is the Overall Equipment Effectiveness (OEE) of the machine?
An operator works an 8.0-hour shift and completes work carrying a standard allowance of 10.0 standard hours. The guaranteed base rate is $30.00 per hour. Under the Halsey 50-50 incentive plan, what are the operator's shift earnings?