11.4 Benchmarking & Performance Attribution
Key Takeaways
- A valid benchmark must be specified in advance, unambiguous, investable, measurable, appropriate to the mandate, and reflective of the manager's current investment opinion.
- Brinson attribution decomposes active return into an allocation effect from sector or asset class weighting decisions, a selection effect from stock picking within each sector, and an interaction effect combining the two.
- The allocation effect for a sector equals the active weight multiplied by the difference between the benchmark sector return and the total benchmark return.
- The selection effect for a sector equals the benchmark weight multiplied by the difference between the portfolio sector return and the benchmark sector return.
- Benchmark choice materially changes the verdict on a manager, so an inappropriate or unrepresentative benchmark can manufacture apparent alpha that is really uncompensated style, credit or currency exposure.
11.4 Benchmarking & Performance Attribution
A return figure on its own tells a client almost nothing. A 9% year is excellent against a benchmark that returned 4% and poor against one that returned 15%. The syllabus therefore requires candidates to understand how benchmarking can be used to measure performance and to know commonly used performance attribution techniques — the two steps that turn a number into a judgement about skill.
What Makes a Benchmark Valid
A benchmark is a passive, investable representation of the mandate the manager was hired to run. Industry practice tests it against seven properties:
| Property | Requirement | What fails it |
|---|---|---|
| Specified in advance | Agreed before the measurement period begins | Choosing the index that flatters the result afterwards ("benchmark shopping") |
| Unambiguous | Constituents and weights are clearly identifiable | "Global equities" with no index named |
| Investable | The manager could passively hold it instead | An index containing unlisted or closed-market securities |
| Measurable | Return can be calculated at the required frequency | An illiquid asset class priced annually against a monthly mandate |
| Appropriate | Consistent with the manager's style and universe | A small-cap value manager measured against a large-cap growth index |
| Reflective of current opinion | The manager has knowledge of, and a view on, the constituents | A UK team benchmarked to an emerging market index |
| Owned / accountable | The manager accepts it as the standard of accountability | A benchmark imposed and then disowned when performance is poor |
Types of benchmark
- Market index. The default for a single-asset-class mandate: FTSE All-Share, MSCI World, a Bloomberg aggregate bond index. Note whether the index is price, total or net return — comparing a portfolio's total return with a benchmark's price return creates fictitious alpha equal to the dividend yield.
- Composite (blended) benchmark. For a multi-asset mandate, a weighted blend such as 60% MSCI World / 40% global aggregate bonds hedged to sterling. The weights should be the client's strategic asset allocation, which makes the composite the single most important benchmark in private wealth management.
- Peer group / universe median. Comparison against other managers. Intuitive for clients but weak analytically: peer universes are not investable, are not specified in advance, and suffer survivorship bias because failed funds are removed from the database.
- Absolute / target return. "Cash plus 4%" or "CPI plus 3%". Appropriate for an unconstrained or liability-aware mandate, but it is not investable and gives no information about whether a poor year was market-wide.
- Liability-based benchmark. The discounted value of the client's future obligations. The most honest benchmark for a decumulating retiree, since it measures whether the plan is on track rather than whether the manager beat an index.
- Custom / strategy benchmark. Built to match the manager's actual universe and constraints where no standard index does.
Performance Attribution
Attribution answers the question where did the active return come from? The standard framework is Brinson attribution, which splits the difference between portfolio and benchmark return into three effects.
For each sector or asset class $i$:
where $w$ is a weight, $R$ a return, subscript $p$ the portfolio, $b$ the benchmark, and $R_b$ the total benchmark return.
- Allocation effect rewards being overweight sectors that beat the overall benchmark and underweight those that lagged. It measures the top-down decision.
- Selection effect rewards picking better-than-index stocks within a sector, holding the weighting neutral. It measures the bottom-up decision.
- Interaction effect captures the combined consequence of overweighting a sector and picking well within it. Many commercial systems fold it into selection; a candidate should know it exists and why.
The three effects sum to the total active return, and a fourth line — currency — is added for an international mandate.
Worked example
A portfolio and its benchmark hold two sectors over one year:
| Sector | Portfolio weight | Benchmark weight | Portfolio return | Benchmark return |
|---|---|---|---|---|
| Technology | 60% | 40% | 14.0% | 12.0% |
| Utilities | 40% | 60% | 5.0% | 6.0% |
Total returns. Portfolio: $(0.60 \times 14%) + (0.40 \times 5%) = 8.4% + 2.0% = 10.4%$ Benchmark: $(0.40 \times 12%) + (0.60 \times 6%) = 4.8% + 3.6% = 8.4%$ Active return = +2.0%
Allocation. Technology: $(0.60 - 0.40) \times (12% - 8.4%) = 0.20 \times 3.6% = +0.72%$ Utilities: $(0.40 - 0.60) \times (6% - 8.4%) = -0.20 \times -2.4% = +0.48%$ Total allocation = +1.20%
Selection. Technology: $0.40 \times (14% - 12%) = +0.80%$ Utilities: $0.60 \times (5% - 6%) = -0.60%$ Total selection = +0.20%
Interaction. Technology: $0.20 \times 2% = +0.40%$ Utilities: $-0.20 \times -1% = +0.20%$ Total interaction = +0.60%
Check: $1.20% + 0.20% + 0.60% = +2.00%$, exactly the active return.
The interpretation matters more than the arithmetic. This manager added most of their value by asset allocation, not stock picking: they were correctly overweight the sector that beat the benchmark and correctly underweight the one that lagged. Their utilities stock selection was actually negative. A client paying a bottom-up stock-picking fee is not getting what they are paying for — and the same result over several periods would justify replacing the equity sleeve with an index fund and paying separately for the allocation call.
Fixed income attribution
Bond attribution decomposes differently, because a bond's return is driven by curve and spread rather than by sector membership. The standard breakdown is duration (interest rate) effect, yield curve positioning, credit / spread effect, sector allocation, security selection, and currency.
Presentation and Standards
- Gross versus net. Attribution is usually run gross of fees; the client experiences net. Both must be shown, and the gap is the fee.
- Composites and GIPS. The Global Investment Performance Standards require firms to group all discretionary, fee-paying portfolios of a similar strategy into a composite, so that a firm cannot advertise only its best account. GIPS compliance is voluntary but is now effectively a condition of institutional business.
- Time period. A single year of attribution is noise. Consistency of the same effect across several periods is what evidences repeatable skill.
- The benchmark decides the verdict. Change the benchmark from a broad market index to a style-matched one and apparent alpha frequently vanishes, because the "skill" was uncompensated exposure to a factor, a credit tier or a currency. Choosing the benchmark honestly is therefore the first act of performance measurement, not an afterthought.
A portfolio holds 30% in Energy against a benchmark weight of 20%. Energy returned 15.0% in the benchmark, while the total benchmark returned 9.0%. What is the allocation effect attributable to the Energy decision?
An investment consultant discovers that a UK equity income manager has been measured against the FTSE 100 total return index, but that the portfolio consistently holds around 40% in mid-cap and small-cap names outside that index. Why is this benchmark problematic?
Why is a peer group universe median generally regarded as an inferior benchmark to a market index, despite being intuitive for clients?