17.1 Historical Cost Databases, Data Organization & Benchmarking
Key Takeaways
- Blueprint tasks 1.A, 1.B, 1.Q, and 1.AA make historical data research, analogous data sets, current-data organization, and historical-data compilation examinable parts of the 36% Managing Project Costs domain.
- A reusable cost record must carry quantity, scope, as-spent date, location, currency, contract type, market conditions, and an actual-versus-estimate flag — a lump sum without context cannot be normalized.
- Normalize in a fixed order: scope, quantity/capacity, escalation by cost-index ratio, location factor, currency, then contract type and market conditions.
- Capacity power-sizing applies to process capacity, not to floor area or linear scope, where unit rates scale close to linearly.
- Ratio benchmarks (engineering as a percentage of installed cost, indirects as a percentage of direct) catch estimate defects that a credible-looking total conceals.
17.1 Historical Cost Databases, Data Organization & Benchmarking
Blueprint task 1.A is "Research historical database" and 1.AA is "Compile historical data." They bracket Domain 1 — the first and last tasks in the largest domain on the exam — because everything else in cost engineering depends on them. An estimate is a prediction built from history; a benchmark is a comparison against history; a forecast is history extrapolated. If the historical database is unusable, every downstream product is guesswork wearing a spreadsheet.
1. What Belongs in a Cost Database
A historical cost record that cannot be reused is just an archive. Reusability requires that each record carry enough context to be normalized onto a new project.
| Data element | Why it must be captured | Failure if omitted |
|---|---|---|
| Quantity and unit of measure | Converts cost to a unit rate | Lump sums cannot be scaled |
| Scope definition | States what was and was not included | Silent scope differences corrupt comparisons |
| Date of the cost (as-spent period) | Anchors the escalation adjustment | Cannot be brought to current dollars |
| Location | Anchors the location factor | Gulf Coast rates applied to remote Arctic work |
| Currency and exchange basis | Enables cross-border reuse | Nominal figures become meaningless |
| Contracting strategy | Lump-sum vs. reimbursable prices behave differently | Contractor risk premium mistaken for base cost |
| Market conditions at award | Bid climate can move price 20%+ independent of scope | Peak-market prices baselined as normal |
| Actual vs. estimate flag | Distinguishes what happened from what was predicted | Estimates recycled as if they were history |
| Basis and exclusions | Records assumptions and carve-outs | Hidden owner-furnished equipment appears free |
[!IMPORTANT] Store actuals, not just final estimates. The most common database defect is a library full of approved estimates rather than as-built costs. An estimate library teaches you what your organization believed; only an actuals library teaches you what things cost.
2. Normalization: The Step That Makes History Comparable
Raw history is never directly reusable. Task 1.B — "develop analogous data sets (e.g., benchmarking, validating estimates, cost comparisons)" — requires normalization first. Apply the adjustments in a fixed order, and record each factor separately so the chain can be audited:
- Scope alignment. Strip or add scope so the two projects describe the same work. This is judgement, not arithmetic, and it is where most of the error lives.
- Quantity/capacity adjustment. Use unit rates for linear scope, or capacity factoring (the power-sizing rule from Section 3.1) for process capacity.
- Time (escalation). Multiply by the ratio of cost indices: Cost(current) = Cost(historical) x (Index_current / Index_historical).
- Location. Apply a location factor for labor productivity, wage rates, freight, duties, and local content requirements.
- Currency. Convert at a stated rate and state the date of that rate.
- Contracting and market. Adjust for contract type and for the bid climate at award.
Worked normalization
A 12,000 m2 warehouse was built in 2021 in the US Gulf Coast for $9,600,000. You need a screening figure for an 18,000 m2 warehouse in a location with a 1.18 location factor. The relevant cost index moved from 228 (2021) to 285 (current).
| Step | Calculation | Result |
|---|---|---|
| Historical unit rate | $9,600,000 / 12,000 m2 | $800 / m2 |
| Escalate to current | $800 x (285 / 228) = $800 x 1.25 | $1,000 / m2 |
| Apply location factor | $1,000 x 1.18 | $1,180 / m2 |
| Apply new quantity | $1,180 x 18,000 m2 | $21,240,000 |
Note what was not done: no capacity exponent was applied, because warehouse area scales close to linearly for the same construction type. Applying a 0.6 exponent here would understate the estimate by roughly 25%. Power-sizing belongs to process capacity, not to floor area.
3. Organizing Current Data (Task 1.Q)
Task 1.Q — "organize current data" — is the discipline that keeps today's project from becoming tomorrow's unusable archive. Three structural decisions do most of the work:
- A stable code of accounts. Costs must be captured against the same code structure from estimate through commitment, accrual, actual, and forecast. If the estimate uses one breakdown and the accounting system another, no variance is ever explainable.
- One transaction, one home. Every cost transaction maps to exactly one control account. Split-coded transactions destroy traceability.
- Cutoff discipline. Every reporting period needs a defined data cutoff, applied identically to earned value, actuals, commitments, and accruals. Mixed cutoffs manufacture variances that do not exist.
4. Benchmarking
Benchmarking compares a project or an estimate against a normalized peer set. It answers "is this number credible?" — a question the estimate's own build-up cannot answer, because a bottom-up estimate can be internally consistent and still wrong.
| Benchmark type | Metric example | Typical use |
|---|---|---|
| Capital intensity | $ per installed kW, $ per bbl/day, $ per m2 | Screening a Class 5 estimate |
| Cost ratio / factor | Bulk materials as % of major equipment; indirects as % of direct | Testing internal proportions |
| Unit rate | Labor-hours per tonne of structural steel erected | Testing craft productivity assumptions |
| Schedule intensity | Months per $100M installed | Testing execution realism |
[!TIP] Ratio benchmarks catch errors that totals hide. If a process plant estimate shows engineering at 4% of total installed cost when the peer set runs 10–14%, the total may still land in a credible range while the engineering allowance is badly wrong. Benchmark the shape of the estimate, not only its magnitude.
Sample-size honesty
Three data points are an anecdote. When you report a benchmark, report the number of data points, the spread, and the normalization applied. A benchmark quoted as a single number, with no range and no basis, invites false precision and is the kind of claim the memo domain expects you to qualify.
A cost engineer needs a screening estimate for a 25,000 m2 distribution warehouse. The historical database contains a 10,000 m2 warehouse of identical construction type built four years ago for $8,000,000. The relevant building cost index has moved from 240 to 288, and the new site carries a location factor of 1.10. What is the screening estimate, and what is the most important modelling choice?
A bottom-up Class 3 estimate for a process plant totals $310 million, which falls comfortably inside the range of the organization's capital-intensity benchmark of $/installed capacity. However, the estimate shows engineering at 4.1% of total installed cost, while the normalized peer set runs 10% to 14%. What conclusion should the cost engineer reach?
Which practice most directly satisfies blueprint task 1.Q, 'organize current data,' so that a live project's cost records remain usable for both variance analysis and future benchmarking?