7.3 Double Lift Charts & Competing Model Evaluation
Key Takeaways
- Double lift charts directly evaluate the relative competitive advantage of a proposed candidate model against an incumbent production model by sorting policies according to the ratio of their predictions.
- The sort ratio (Predicted Loss Cost Candidate / Predicted Loss Cost Incumbent) isolates segments of maximum model disagreement; extreme deciles represent policies where one model predicts heavy loss cost while the other predicts low risk.
- A candidate model demonstrates dominance over the incumbent if the actual holdout loss costs strictly follow the candidate model's predictions across all deciles, proving that the incumbent's pricing errors reflect genuine underwriting risk.
- In an adverse selection context, policies in Decile 10 (where the candidate predicts much higher loss costs than the incumbent) represent severe underpricing vulnerabilities under current rates, while Decile 1 identifies overcharged risks vulnerable to competitor poaching.
- Double lift charts must be constructed on independent out-of-sample or out-of-time data using exposure weighting; fitting on training data artificially guarantees an unearned advantage for more parameterized models.
7.3 Double Lift Charts & Competing Model Evaluation
Exam Focus: In property and casualty ratemaking, an actuary rarely builds a predictive model in a vacuum. Instead, the task is almost always to determine whether a newly developed candidate model (Model B, such as a refit or expanded GLM) should replace an existing production model (Model A, such as the current rating tariff or legacy GLM). Standard single lift charts often fail in this task because major rating factors (e.g., driver age or territory) cause both models to exhibit steep single lift curves. The Double Lift Chart (also known as a Two-Model Comparison Lift Chart or Displacement Chart) directly isolates the disagreement between the two models to prove which model captures true underlying risk. Candidates must master its construction methodology, the sort ratio interpretation, and how to identify adverse selection vulnerabilities.
When two models are evaluated using standard metrics, Model B may show a slightly lower deviance or higher Gini than Model A. However, corporate executives and chief actuaries must weigh the substantial commercial disruption of re-filing rates, reprogramming policy administration engines, and disrupting existing policyholder premiums. The double lift chart provides the definitive economic proof required to justify model replacement.
1. Why Single Lift Charts Are Insufficient for Model Replacement
To understand why double lift charts are necessary, consider evaluating two models on an independent holdout dataset of personal auto policies:
- Model A (Incumbent Rating Plan): A traditional Generalized Linear Model with 15 rating variables.
- Model B (Candidate Model): A refit GLM with 28 rating variables, adding banded driver age, a territory-by-vehicle-class interaction, and credibility-grouped construction codes.
If the actuary constructs a single lift chart for Model A and a single lift chart for Model B, both charts will look remarkably similar. Both models will show a steep, monotonic progression from Decile 1 to Decile 10, achieving similar top-to-bottom lift ratios (e.g., 5.8x vs. 6.2x). Why?
Because fundamental insurance variables (such as driver age, prior at-fault accidents, vehicle age, and basic territory) dominate the aggregate distribution of losses. Both models identify that an 18-year-old with two prior accidents is much riskier than a 50-year-old with a clean record. Single lift charts mask the marginal improvements of Model B because the dominant shared signals drown out the nuanced differences.
THE COMPARISON DILEMMA
Model A (Incumbent GLM) Model B (Candidate GBDT)
Single Lift: Steep & Monotonic (6.0x) Single Lift: Steep & Monotonic (6.3x)
▲ ▲
│ ┌─────┘ │ ┌─────┘
│ ┌─────┘ │ ┌─────┘
└──┴────────────────► └──┴────────────────►
QUESTION: Does Model B's incremental 0.3x lift justify a multi-million-dollar
rating overhaul, or is it just fitting noise on shared factors?
SOLUTION: The Double Lift Chart isolates the DISAGREEMENT between Model A and Model B!
The double lift chart strips away the common ground and directly evaluates the policies where Model A and Model B make divergent predictions.
2. Construction Methodology of the Double Lift Chart
The double lift chart is constructed on an independent holdout test dataset or out-of-time validation dataset via a rigorous six-step protocol.
DOUBLE LIFT CHART CONSTRUCTION WORKFLOW
┌────────────────────────────────────────────────────────────────────────┐
│ Step 1: Generate Predictions for Both Models on Holdout Data │
│ For each policy i, compute y_hat_A_i (Incumbent) and y_hat_B_i (Candidate)│
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Step 2: Calculate the Sort Ratio (Relative Model Relativity) │
│ Compute Ratio_i = y_hat_B_i / y_hat_A_i │
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Step 3: Sort Policies Ascending by the Sort Ratio │
│ Rank policies: Ratio_(1) <= Ratio_(2) <= ... <= Ratio_(n) │
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Step 4: Partition into Quantiles of EQUAL EARNED EXPOSURE │
│ Accumulate earned exposure w_i into 10 deciles of equal exposure │
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Step 5: Understand the Physical Meaning of the Extreme Deciles │
│ Decile 1: Model B discounts vs. Model A (Candidate thinks safer) │
│ Decile 10: Model B surcharges vs. Model A (Candidate thinks riskier) │
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Step 6: Compute & Plot Decile Metrics │
│ For each decile, plot: Actual Loss Cost vs. Model A vs. Model B │
└────────────────────────────────────────────────────────────────────────┘
Step 1: Generate Predictions
For every policy $i$ in the holdout partition, obtain the predicted loss cost from Model A ($\hat{y}{A,i}$) and Model B ($\hat{y}{B,i}$).
Step 2: Compute the Sort Ratio
For each holdout policy, calculate the Sort Ratio ($R_i$), defined as the ratio of the candidate model prediction to the incumbent model prediction:
(Note: Alternatively, actuaries may define the ratio on a logarithmic scale: $\Delta_i = \ln(\hat{y}{B,i}) - \ln(\hat{y}{A,i})$. Both methods produce identical policy rank-orderings).
Step 3: Sort Policies Ascending by Sort Ratio
Sort all holdout policies in strictly ascending order of $R_i$:
Step 4: Partition into Equal Earned Exposure Deciles
Just as with single lift charts, divide the sorted portfolio into 10 deciles $\mathcal{Q}k$ ($k=1, \dots, 10$) of equal earned exposure $W_k = \sum{i \in \mathcal{Q}k} w_i = W{\text{total}} / 10$.
Step 5: Actuarial Interpretation of the Quantile Spectrum
Understanding what each decile physically represents in the insurance market is essential for interpreting the results:
Decile 1 Decile 3 Deciles 5-6 Decile 8 Decile 10
Ratio << 1.0 Ratio < 1.0 Ratio ≈ 1.0 Ratio > 1.0 Ratio >> 1.0
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Model B │ │ Model B │ │ AGREEMENT │ │ Model B │ │ Model B │
│ STRONGLY │ │ Moderately │ │ ZONE │ │ Moderately │ │ STRONGLY │
│ DISCOUNTS │ │ Discounts │ │ y_B ≈ y_A │ │ Surcharges │ │ SURCHARGES │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘
Model B says risk Model B prices Both models agree Model B prices Model B says risk
is vastly lower cheaper than on expected loss higher than is vastly higher
than Model A says. Model A. cost. Model A. than Model A says.
- Decile 1 (Lowest Ratio, $R_i \ll 1.0$): Policies where Model B predicts a much lower loss cost than Model A. (Example: $R_i = 0.55 \implies$ Model B believes the risk costs $550, while Model A prices it at $1,000).
- Deciles 5 & 6 (Intermediate Ratios, $R_i \approx 1.0$): Policies where Model A and Model B agree on the price.
- Decile 10 (Highest Ratio, $R_i \gg 1.0$): Policies where Model B predicts a much higher loss cost than Model A. (Example: $R_i = 1.85 \implies$ Model B believes the risk costs $1,850, while Model A prices it at $1,000).
Step 6: Compute and Plot Decile Metrics
For each decile $k$, compute three exposure-weighted metrics:
Plot all three series across the 10 deciles on the same axes.
3. Interpreting Model Dominance
The fundamental diagnostic question answered by the double lift chart is: When Model A and Model B disagree, who is right? The answer is revealed by whether the Actual Loss Cost line ($\bar{Y}_k$) follows Model B or Model A.
CASE 1: MODEL B DOMINATES MODEL A
Loss Cost
│
│ ▲ Actual Loss Cost (Tracks B!)
│ ┌─────┘
│ ┌─────┘ ─── Model B Predicted Loss Cost
│ ┌─────┘
│ ┌─────┘
│ ┌─────────────────────.─────┴───────────────────────┐ Model A Predicted Loss Cost
│ │ .───── │ (Relatively flat across sort ratio)
│ │ .───── │
│ │ .───── │
│ └─── │
│
└────────────────────────────────────────────────────────►
D1 D2 D3 D4 D5 D6 D7 D8 D9 D10
Deciles of Sort Ratio (y_hat_B / y_hat_A)
Case 1: Model B Dominates Model A (Clear Winning Candidate)
- Visual Shape: The Actual Loss Cost curve slopes steeply upward, mirroring the steep slope of Model B's predictions, while Model A's predicted loss cost line remains relatively flat across the deciles.
- Actuarial Interpretation: In Decile 1, where Model B predicted low losses, actual losses were indeed low (confirming Model A was overpriced). In Decile 10, where Model B predicted high losses, actual losses were indeed high (confirming Model A was dangerously underpriced). Model B's disagreements with Model A represent genuine actuarial signal. Model B should be approved for production.
Case 2: Model A Retains Dominance (Candidate Model Overfitting)
- Visual Shape: The Actual Loss Cost line remains flat, closely hugging Model A's flat prediction line, while Model B's steep upward predictions are completely contradicted by empirical reality.
- Actuarial Interpretation: When Model B surcharged policies in Decile 10, those policies actually incurred normal losses. When Model B discounted policies in Decile 1, they incurred normal losses. Model B's departures from Model A were driven by sample variance and over-fitting. Model B must be rejected.
Diagnostic Summary Table
| Double Lift Shape | Behavior of Actual Loss Cost Line | Actuarial Verdict | Underwriting Decision |
|---|---|---|---|
| Model B Dominates | Steep upward slope; closely hugs Model B across all 10 deciles. | Model B captures true risk differentiation missed by Model A. | Adopt Model B; re-file rating manual; adjust underwriting tiers. |
| Model A Dominates | Horizontal line; hugs Model A; ignores Model B's slope. | Model B's added complexity is fitting pure sample noise. | Retain Model A; reject Model B. |
| Symmetric Non-Dominance | Actual line sits halfway between Model A and Model B. | Both models possess unique, non-overlapping predictive signal. | Create an ensemble/blended model: $\hat{y} = 0.5\hat{y}_A + 0.5\hat{y}_B$. |
| Tail Divergence Only | Deciles 2–9 are flat, but Decile 10 shoots upward. | Model B identified a specific high-hazard underwriting rule or shock niche. | Incorporate Model B's tail features into Model A as underwriting filters. |
4. The Loss Ratio Double Lift Chart Variant
In property and casualty companies where rates in the holdout period were charged according to Model A, actuaries construct a highly intuitive operational variant: the Loss Ratio Double Lift Chart.
Instead of plotting loss costs, compute the Actual Loss Ratio under Incumbent Premium for each decile:
Actual Loss Ratio (under Model A Rates)
│
140% │ ▲ Decile 10: 142% Loss Ratio
│ ┌─────┘ (Underwriting Catastrophe!)
120% │ ┌─────┘
│ ┌─────┘
100% │ ┌───┘
│ ┌──────┘
80% │═══════════════════════════════════════════════════════ Target Loss Ratio = 70%
│ ┌─────┘
60% │ ┌─────┘
│ ┌─────┘
40% │┌─────┘ Decile 1: 38% Loss Ratio
│┘ (Highly Profitable / Poaching Risk)
0% └────────────────────────────────────────────────────────►
D1 D2 D3 D4 D5 D6 D7 D8 D9 D10
Deciles of Sort Ratio (y_hat_B / y_hat_A)
Actuarial Interpretation of the Loss Ratio Variant
- If Model A were perfect, the actual loss ratio would be completely flat across all deciles (e.g., exactly 70% in every decile), because Model A's premiums would perfectly match expected loss costs.
- When Model B is superior, the loss ratio curve under Model A rates exhibits a violent upward slope:
- Decile 1: Has an actual loss ratio of 38%. These policies are wildly profitable under Model A's rates because Model A is severely overcharging them.
- Decile 10: Has an actual loss ratio of 142%. These policies are an underwriting disaster under Model A's rates because Model A is severely undercharging them.
5. Adverse Selection Dynamics & Commercial Strategy
The double lift chart is not merely a statistical diagnostic; it is a simulation of competitive market dynamics.
ADVERSE SELECTION SPIRAL UNDER MODEL A
┌─────────────────────────────────────────────────────────────────┐
│ Competitors implement Model B while our company stays on Model A│
└────────────────────────────────┬────────────────────────────────┘
│
┌──────────────────────┴──────────────────────┐
▼ ▼
┌─────────────────────┐ ┌─────────────────────┐
│ Decile 1 │ │ Decile 10 │
│ Preferred Risks │ │ Hazardous Risks │
└──────────┬──────────┘ └──────────┬──────────┘
│ │
▼ ▼
Competitor offers 30% Competitor surcharges +50%
discount based on Model B. or cancels based on Model B.
│ │
▼ ▼
Our good customers LEAVE. Hazardous risks FLOCK to us
Carrier loses profitable because our Model A rates
exposure volume. are far too cheap!
│ │
└──────────────────────┬──────────────────────┘
│
▼
Carrier's overall Loss Ratio explodes;
Average rates must rise, fueling the death spiral!
The Operational Underwriting Action Plan
- Capitalizing on Decile 1 (Growth Opportunity):
- Under current Model A pricing, the insurer is overcharging Decile 1 accounts. Competitors with advanced models will inevitably identify these accounts and poach them with lower quotes.
- By implementing Model B, the insurer can proactively reduce rates on Decile 1, improving customer retention and aggressively marketing to acquire high-margin, preferred risks.
- Defending Against Decile 10 (Adverse Selection Defense):
- Under current Model A pricing, Decile 10 policies generate a 140%+ loss ratio. The insurer is acting as a 'magnet' for risks that competitors reject or heavily surcharge.
- Implementing Model B allows the insurer to impose necessary rate surcharges, restrict underwriting guidelines (e.g., requiring senior underwriter sign-off or telematics enrollment), or non-renew severely unprofitable accounts.
6. Actuarial Traps & Exam Pitfalls
┌────────────────────────────────────────────────────────┐
│ Actuarial Traps in Double Lift Charts │
└────────────────────────────────────────────────────────┘
│
┌───────────────────────────────────────┼───────────────────────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Inverting │ │ Training Data│ │ Ignoring │
│ Sort Ratio │ │ Evaluation │ │ Credibility │
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
▼ ▼ ▼
Reversing y_hat_B / y_hat_A Running double lift on Small exposure in
inverts deciles; causes training set gives false outer deciles produces
analyst to confuse over- guarantee of candidate erratic loss spikes;
priced vs underpriced tiers. model dominance. must enforce equal exposure.
Trap 1: Inverting the Sort Ratio Formula
Candidates often invert the ratio, calculating $R_i = \hat{y}{A,i} / \hat{y}{B,i}$ instead of $\hat{y}{B,i} / \hat{y}{A,i}$. If the ratio is inverted, the entire decile spectrum flips: Decile 1 becomes the tier where Model B surcharges, and Decile 10 becomes the tier where Model B discounts. Always confirm the ratio definition before interpreting whether Decile 1 represents underpriced or overpriced business.
Trap 2: Running Double Lift on Training Data
A complex machine learning model (such as a 500-tree GBDT) evaluated against a simple GLM on the training dataset will always dominate the double lift chart. Its second-order optimization allows it to fit training sample residuals, producing an artificial appearance of superiority. A double lift chart is valid only when constructed on independent holdout or out-of-time validation data.
Trap 3: Sorting Without Exposure Weighting
If the double lift chart quantiles are formed using raw policy count rather than earned exposure, high sort ratio deciles can become dominated by volatile, low-exposure policies (such as a 3-day policy with a high predicted rate). This introduces massive sampling variance in Decile 1 and Decile 10, obscuring true model performance.
In a double lift chart comparing an incumbent rating plan (Model A) against a candidate refit GLM (Model B), policies in the holdout dataset are sorted ascending by the ratio y_hat_B / y_hat_A and partitioned into equal-exposure deciles. If actual holdout loss costs increase steeply from Decile 1 to Decile 10, closely tracking Model B's predictions, what does this demonstrate?
An actuary constructs a double lift chart sorting policies by the ratio of Candidate Model B to Current Model A (y_hat_B / y_hat_A). The current book of business is priced using Model A rates. What operational risk exists for policies falling into Decile 10 if the carrier delays implementing Model B?
What is the primary reason why P&C pricing actuaries construct double lift charts rather than relying solely on single lift charts when evaluating whether to replace an existing rating model?