14.2 Predictive Analytics & Machine Learning in Underwriting

Key Takeaways

  • Predictive modeling replaces traditional univariate tabular ratemaking with multivariate statistical algorithms that evaluate dozens of risk factors simultaneously, eliminating omitted variable bias and cross-subsidies.

  • Generalized Linear Models (GLMs) are the foundational industry standard for ratemaking, utilizing Poisson distributions with exposure offset terms to model claim frequency, and Gamma or Lognormal distributions to model claim severity.

  • Advanced machine learning algorithms—including Random Forests and Gradient Boosted Machines (XGBoost, LightGBM)—capture complex non-linearities and variable interactions, delivering superior predictive accuracy over traditional linear structures.

  • Underwriting operations leverage predictive scorecards to execute automated triage, routing clean submissions through Straight-Through Processing (STP) while flagging complex or high-hazard risks for human underwriter evaluation.

  • Credit-based insurance scores exhibit strong actuarial correlation with loss frequency and severity, but their deployment must strictly comply with the Fair Credit Reporting Act (FCRA) and diverse state regulatory prohibitions.

Last updated: September 2026

Predictive Analytics & Machine Learning in Underwriting

Quick Answer: Underwriting predictive analytics uses statistical and machine learning algorithms to estimate an applicant's expected loss cost and loss frequency. The insurance industry's foundational standard is the Generalized Linear Model (GLM), which separates risk into Claim Frequency (modeled via Poisson distribution with an exposure offset term) and Claim Severity (modeled via Gamma or Lognormal distribution). Modern carriers supplement GLMs with advanced machine learning like Gradient Boosted Machines (GBMs) and Random Forests to capture complex interactions. Key operational applications include Straight-Through Processing (STP) for automated underwriting triage and Credit-Based Insurance Scores (CBIS) under strict FCRA compliance. Models are evaluated using Lift Charts, Gini coefficients, ROC/AUC curves, and confusion matrices.


1. The Evolution of Actuarial Rating: From Tabular Plans to Multivariate Models

Historically, property-casualty ratemaking relied upon univariate (one-way) tabular rating plans. Actuaries analyzed loss ratios across a single variable at a time (e.g., territory or vehicle type) while holding all other factors conceptually constant.

However, one-way analysis suffers from severe statistical distortions:

  • Omitted Variable Bias: Fails to control for confounding variables correlated with both the analyzed factor and claim losses.
  • Distortion from Correlated Factors: If youthful drivers disproportionately drive older sports cars in urban territories, a univariate analysis will double-count risk across age, vehicle class, and geography.
  • Inefficient Cross-Subsidies: Good risks within high-risk categories are overcharged, while poor risks in low-risk categories are undercharged, triggering severe adverse selection against the carrier.

To solve this, modern insurance pricing relies on multivariate predictive modeling, evaluating all rating factors simultaneously to isolate each variable's true independent marginal effect.

                         THE ACTUARIAL EVOLUTION

   UNIVARIATE TABULAR RATING                   MULTIVARIATE PREDICTIVE RATING
┌───────────────────────────────┐            ┌───────────────────────────────┐
│ • One variable analyzed alone │            │ • All variables modeled jointly│
│ • Confounded risk factors     │    ───►    │ • Confounders isolated        │
│ • Inherent cross-subsidies    │            │ • Multiplicative relativities │
│ • Vulnerable to adverse select│            │ • Accurate risk-cost alignment│
└───────────────────────────────┘            └───────────────────────────────┘

2. Generalized Linear Models (GLMs): The Actuarial Gold Standard

Generalized Linear Models (GLMs) represent the foundational statistical framework for property-casualty ratemaking. Introduced by Nelder and Wedderburn, GLMs extend classical ordinary least squares (OLS) regression to accommodate insurance data, which violates standard normality and homoscedasticity assumptions.

The Mathematical Structure of a GLM

A GLM consists of three core components:

  1. Random Component: The probability distribution of the dependent response variable Y, chosen from the exponential dispersion family (e.g., Poisson, Gamma, Binomial).
  2. Systematic Component: A linear combination of independent explanatory variables (rating factors): η=β0+β1X1+β2X2+⋯+βkXk=Xβ\eta = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \dots + \beta_k X_k = X\beta
  3. Link Function g(·): A monotonic mathematical function that connects the expected value of the response variable μ = E(Y) to the systematic linear predictor: η=g(μ)⟺μ=g−1(Xβ)\eta = g(\mu) \quad \Longleftrightarrow \quad \mu = g^{-1}(X\beta)

Why Ordinary Least Squares (OLS) Fails for Insurance Losses

Under standard OLS regression, the response variable is assumed to follow a continuous normal (bell-shaped) distribution with constant variance. Insurance claims violate these assumptions fundamentally:

  • Claim Frequency consists of non-negative integers (0, 1, 2, …), with the vast majority of policyholders experiencing zero claims.
  • Claim Severity (loss dollar amount per claim) is strictly positive, non-negative, and exhibits extreme positive (right-hand) skewness with heavy tails representing catastrophic payouts.
  • OLS linear regression would predict negative claim counts and negative loss dollar amounts for low-risk policyholders, which is physically and mathematically impossible.

Deconstructing Pure Premium: Two-Part GLM Modeling

In property-casualty actuarial practice, the Pure Premium (also known as the Loss Cost—the expected loss dollar amount per exposure unit) is decomposed into two distinct, independently estimated GLMs:

Pure Premium=E(Frequency)×E(Severity)\text{Pure Premium} = E(\text{Frequency}) \times E(\text{Severity}) Pure Premium=(Number of ClaimsExposures)×(Total Dollar LossesNumber of Claims)=Total Dollar LossesExposures\text{Pure Premium} = \left(\frac{\text{Number of Claims}}{\text{Exposures}}\right) \times \left(\frac{\text{Total Dollar Losses}}{\text{Number of Claims}}\right) = \frac{\text{Total Dollar Losses}}{\text{Exposures}}
                       TWO-PART GLM ARCHITECTURE
                                   │
         ┌─────────────────────────┴─────────────────────────┐
         ▼                                                   ▼
 [FREQUENCY MODEL]                                   [SEVERITY MODEL]
  • Distribution: Poisson or Negative Binomial        • Distribution: Gamma or Lognormal
  • Target: Claim Counts                              • Target: Paid / Incurred Loss per Claim
  • Link: Log Link ln(μ)                              • Link: Log Link ln(μ)
  • Exposure Offset: ln(Exposure)                     • Weight: Claim Count per Record
         │                                                   │
         └─────────────────────────┬─────────────────────────┘
                                   ▼
                  [ PURE PREMIUM (LOSS COST) ]
                  Expected Frequency × Expected Severity

A. Frequency Modeling: The Poisson Distribution and the Exposure Offset

  • Distribution Choice: Claim counts are modeled using the Poisson distribution, or the Negative Binomial distribution when significant overdispersion exists (where variance exceeds the mean).
  • Link Function: The Log Link Function, g(μ) = ln(μ), ensuring that expected frequency is always strictly positive (e^(Xβ) > 0).
  • The Exposure Offset Term: A critical actuarial concept tested in CPCU 550 is the Offset Term. Claim frequency naturally scales with the length of exposure (e.g., a commercial vehicle insured for 12 months has twice the exposure of one insured for 6 months): ln⁡(μExposure)=Xβ  ⟹  ln⁡(μ)=Xβ+ln⁡(Exposure)\ln\left(\frac{\mu}{\text{Exposure}}\right) = X\beta \quad \implies \quad \ln(\mu) = X\beta + \ln(\text{Exposure}) By including ln(Exposure) as an offset term with a regression coefficient fixed strictly at 1.0, the model normalizes frequency across differing policy exposure durations.

B. Severity Modeling: The Gamma Distribution

  • Distribution Choice: Average claim severity is modeled using the Gamma distribution (or Lognormal distribution). The Gamma distribution is defined strictly on positive real numbers (Y > 0) and its variance increases proportionally with the square of the mean, perfectly mirroring empirical claims behavior where larger claims exhibit greater dollar volatility.
  • Link Function: The Log Link Function, g(μ) = ln(μ).

C. Multiplicative Rating Relativities

Because both the frequency and severity models utilize log link functions, converting the predictions back to the natural dollar scale via the inverse link (g⁻¹(η) = e^(η)) results in a multiplicative rating structure:

Pure Premium=eβ0×eβ1X1×eβ2X2×⋯×eβkXk\text{Pure Premium} = e^{\beta_0} \times e^{\beta_1 X_1} \times e^{\beta_2 X_2} \times \dots \times e^{\beta_k X_k} Rate=Base Rate×Factor1×Factor2×⋯×Factork\text{Rate} = \text{Base Rate} \times \text{Factor}_1 \times \text{Factor}_2 \times \dots \times \text{Factor}_k

This multiplicative format is universally favored by state insurance regulators and underwriting rating engines because each factor acts as an intuitive percentage surcharge or discount relativity (e.g., youthful driver factor = 1.45; anti-theft device factor = 0.92).


3. Modern Machine Learning Algorithms in Underwriting

While GLMs remain dominant for regulatory rate filings, insurers increasingly deploy advanced machine learning (ML) algorithms for risk segmentation, tiering, and real-time underwriting triage.

                      MACHINE LEARNING PARADIGMS
                                   │
         ┌─────────────────────────┴─────────────────────────┐
         ▼                                                   ▼
 [SUPERVISED LEARNING]                               [UNSUPERVISED LEARNING]
  • Labeled historical target outcomes                • Unlabeled data exploration
  • Target: Loss cost, renewal lapse, claim fraud     • Target: Latent risk clusters, anomalies
  • Algorithms: GLMs, Trees, Random Forests, GBMs     • Algorithms: K-Means, PCA, Autoencoders

Comparison of Core Machine Learning Algorithms

AlgorithmOperational MechanismKey Underwriting AdvantagesLimitations & Regulatory Challenges
Decision Trees (CART)Recursively partitions data into binary subsets based on splitting criteria (Gini impurity, variance reduction).Highly visual and transparent; easy to translate into manual underwriting guidelines.Highly prone to overfitting; high variance; unstable to minor changes in training data.
Random ForestsAn ensemble bagging (bootstrap aggregation) method that trains hundreds of de-correlated decision trees on random data subsets and random feature selections, averaging their outputs.Drastically reduces model variance; handles high-dimensional data; resistant to overfitting; captures non-linear relationships.Computationally intensive; complex "black-box" structure that cannot be easily expressed as a simple multiplicative rating table.
Gradient Boosted Machines (GBMs) (XGBoost, LightGBM, CatBoost)An ensemble boosting method that builds trees sequentially; each new tree is specifically trained to predict the residual errors (gradients) of the preceding trees.Exceptional predictive precision; current industry gold standard for complex pricing tiering and risk scoring; handles missing values seamlessly.Highly sensitive to hyperparameter tuning (can overfit if learning rate is too high or tree depth too large); requires explainability tools (SHAP values) for regulatory approval.
Neural Networks & Deep LearningInterconnected layers of artificial neurons with non-linear activation functions processing complex multi-layer feature embeddings.Unrivaled performance on unstructured inputs (computer vision photo appraisal, audio transcripts, NLP text).Extreme "black-box" opacity; nearly impossible to justify to state insurance departments for baseline personal lines pricing.

4. Underwriting Applications: Automated Triage, Tiering, and Credit Scores

Automated Underwriting Triage & Straight-Through Processing (STP)

Traditional underwriting required human review of every commercial or personal submission. Modern carriers implement automated predictive triage engines at the point of sale:

                   AUTOMATED UNDERWRITING TRIAGE FUNNEL

                        [Incoming Submissions]
                                   │
                                   ▼
                   [Predictive Scoring Engine]
           (Loss Propensity Score + Appetite Rules Engine)
                                   │
       ┌───────────────────────────┼───────────────────────────┐
       ▼                           ▼                           ▼
[Straight-Through]          [Referral Queue]           [Automated Decline]
  (Score: 0 - 30)            (Score: 31 - 75)            (Score: 76 - 100)
• Clean, low-hazard risks   • Borderline hazard risks  • Risks outside appetite
• Instant electronic bind   • Complex liability issues • Severe loss history
• Zero human intervention   • Routed to human underwriter• Instant electronic rejection
  • Straight-Through Processing (STP): Submissions with high predictive confidence, low predicted loss costs, and strict compliance with automated guidelines are automatically bound electronically within seconds, drastically cutting operational acquisition expense.
  • Human Referral Queue: Submissions with ambiguous data, marginal scores, large exposure limits, or complex loss histories are flagged with specific guidance and routed directly to experienced line underwriters.

Tiered Pricing Algorithms

Rather than sorting insureds into 3 or 4 broad rating classes (e.g., Preferred, Standard, Non-Standard), predictive algorithms segment risks into dozens of granular pricing tiers (e.g., Tier 1 through Tier 50). This micro-segmentation prevents competitors from "cherry-picking" the best risks out of a broad standard class, stabilizing the carrier's combined ratio.

Credit-Based Insurance Scores (CBIS)

A Credit-Based Insurance Score is a numerical index derived from a consumer's credit report data (payment history, outstanding debt balances, length of credit history, pursuit of new credit, and credit mix).

                      CREDIT-BASED INSURANCE SCORING

   Credit Report Elements                 Statistical Fact               Regulatory Framework
┌────────────────────────────┐         ┌───────────────────────┐       ┌──────────────────────┐
│ • Debt-to-credit ratio     │         │ Actuarial studies show│       │ Fair Credit Reporting│
│ • Delinquency records      │  ─────► │ strong correlation to │ ────► │ Act (FCRA) Compliance│
│ • Credit history length    │         │ claim frequency and   │       │ Adverse Action Notice│
│ • Inquiries & new accounts │         │ loss severity         │       │ State Law Prohibitions
└────────────────────────────┘         └───────────────────────┘       └──────────────────────┘
  • Actuarial Validity: Multiple independent actuarial and FTC studies have verified that credit-based insurance scores exhibit a powerful, statistically significant correlation with insurance loss frequency and severity. Individuals who manage their financial credit responsibly demonstrate similar risk-avoidance behaviors behind the wheel or in home maintenance.
  • Insurance Score vs. Financial Credit Score: An insurance score is not a financial credit score (such as a lending FICO score). Insurance scorecards exclude consumer income, asset balances, net worth, and demographic data, focusing strictly on credit management stability.
  • Regulatory Compliance & FCRA: Under the federal Fair Credit Reporting Act (FCRA), if an insurer charges a higher premium, denies coverage, or moves an applicant to a less favorable tier based even partially on a credit report, the carrier must provide an Adverse Action Notice. This notice must inform the consumer of the credit bureau used, explain that the bureau did not make the underwriting decision, and provide the primary risk factors (e.g., "high revolving balances relative to credit limits") that negatively impacted their score.
  • State Prohibitions: Several states (e.g., California, Massachusetts, Maryland, Washington) strictly prohibit or heavily restrict the use of credit-based insurance scores in personal lines ratemaking, citing concerns over socio-economic disparate impact.

5. Model Evaluation Metrics & Validation Frameworks

An algorithm cannot be deployed into production without rigorous validation across multiple statistical dimensions.

                      MODEL VALIDATION SCORECARD
                                   │
     ┌─────────────────────────────┼─────────────────────────────┐
     ▼                             ▼                             ▼
[Lift Charts / Gini]        [ROC Curve / AUC]            [Confusion Matrix]
 • Ranks risks into deciles   • True Pos vs False Pos      • Classification accuracy
 • Measures pricing separation• Evaluates binary triage    • Precision, Recall, F1
 • Top-to-bottom quantile ratio• Target: AUC > 0.75-0.80   • Type I vs Type II errors

1. Lift Charts and the Gini Coefficient

  • Lift Chart (Quantile Plot): Sorts all policies into deciles (10 equal groups of 10%) ranked from lowest predicted risk (Decile 1) to highest predicted risk (Decile 10). The chart plots the actual observed loss cost for each decile. A steep, monotonic upward slope indicates a highly effective model that successfully separates superior risks from inferior risks. Model Lift is calculated as the ratio of actual losses in the worst decile to actual losses in the best decile: Lift=Actual Loss Cost in Decile 10Actual Loss Cost in Decile 1\text{Lift} = \frac{\text{Actual Loss Cost in Decile 10}}{\text{Actual Loss Cost in Decile 1}}
  • Gini Coefficient: Derived from the Lorenz curve, the Gini coefficient measures the area between the model's cumulative loss distribution curve and the 45-degree diagonal line of equal distribution. A Gini of 0.0 indicates no predictive ability, while higher values (e.g., 0.35–0.55) indicate outstanding risk rank-ordering.

2. ROC Curves and Area Under the Curve (AUC)

For binary classification models (e.g., bound vs. not bound, fraudulent vs. non-fraudulent, litigated vs. non-litigated), performance is plotted via the Receiver Operating Characteristic (ROC) curve:

  • Y-Axis: True Positive Rate (Sensitivity / Recall) = TP / (TP + FN)
  • X-Axis: False Positive Rate (1 - Specificity) = FP / (FP + TN)
  • Area Under the Curve (AUC): Measures aggregate discriminatory power. An AUC of 0.50 represents random guessing, while an AUC of 1.00 represents flawless separation. High-performing underwriting triage models typically achieve AUCs between 0.75 and 0.88.

3. The Confusion Matrix and Error Types

In underwriting triage, predictions fall into a 2 × 2 matrix:

Actually High Risk (Referral)Actually Low Risk (STP)
Model Predicts High RiskTrue Positive (TP): Correctly sent to human underwriter.False Positive (FP) [Type I Error]: Clean risk needlessly referred to human, wasting underwriting expense.
Model Predicts Low RiskFalse Negative (FN) [Type II Error]: Hazardous risk mistakenly auto-bound via STP, causing underwriting loss.True Negative (TN): Clean risk properly auto-bound via STP.

Important

Underwriting Asymmetry: In insurance underwriting, a False Negative (Type II Error) is far more economically dangerous than a False Positive (Type I Error). Auto-binding a catastrophic loss exposure destroys capital, whereas needlessly reviewing a clean file merely consumes modest administrative time.

4. Overfitting vs. Underfitting and Cross-Validation

  • Overfitting: Occurs when a complex machine learning model (e.g., a deep decision tree with 20 levels) memorizes the random noise and idiosyncrasies of historical training data. The model demonstrates near-zero error on training records but collapses completely when exposed to unseen future policies.
  • Validation Safeguards:
    • Data Partitioning: Splitting data into Training Set (60%), Validation Set (20%), and Test Holdout Set (20%).
    • Out-of-Time Testing: Testing the model on a future policy year not included in the training window (e.g., train on 2022–2024, test on 2025–2026) to verify stability across economic cycles.
    • k-Fold Cross-Validation: Dividing the training data into k equal subsets, training on k-1 subsets, and validating on the remaining subset, repeated k times.
    • Regularization: Applying L₁ Lasso (shrinking irrelevant coefficients to zero) or L₂ Ridge regression penalties to constrain model complexity.

6. Comprehensive Worked Mathematical Scenario: Fleet Delivery Pure Premium Modeling

The Operational Setup

Pinnacle Commercial Assurance is calibrating a two-part GLM to price commercial last-mile delivery vans. The actuarial team builds independent Poisson frequency and Gamma severity models using earned vehicle years as the exposure base.

Model Parameters & Coefficients

A. Frequency GLM (Poisson Distribution, Log Link, Exposure Offset)

ln⁡(μFreq)=ln⁡(Exposure)+β0+βRadius+βTelematics+βSafety\ln(\mu_{\text{Freq}}) = \ln(\text{Exposure}) + \beta_0 + \beta_{\text{Radius}} + \beta_{\text{Telematics}} + \beta_{\text{Safety}}
  • Base Intercept (β₀): -1.6094 (e^(-1.6094) = 0.2000 base frequency)
  • Operating Radius > 50 miles (β(Radius)): +0.2624 (e^0.2624 = 1.30 relativity)
  • Telematics PHYD Installed (β(Telematics)): -0.2231 (e^(-0.2231) = 0.80 relativity)
  • Advanced Driver Assistance (ADAS) (β(Safety)): -0.1054 (e^(-0.1054) = 0.90 relativity)

B. Severity GLM (Gamma Distribution, Log Link)

ln⁡(μSev)=α0+αRadius+αVehicle Type\ln(\mu_{\text{Sev}}) = \alpha_0 + \alpha_{\text{Radius}} + \alpha_{\text{Vehicle Type}}
  • Base Intercept (α₀): 8.5172 (e^8.5172 = $5,000 base severity)
  • Operating Radius > 50 miles (α(Radius)): +0.1823 (e^0.1823 = 1.20 relativity)
  • Heavy Cargo Chassis (α(Vehicle Type)): +0.2231 (e^0.2231 = 1.25 relativity)

Step-by-Step Calculation for a Commercial Fleet Account

Consider an applicant operating a fleet of 50 delivery vans over a 1.0-year policy term (50 vehicle years of exposure). The fleet operates on an extended radius (> 50 miles), has installed telematics PHYD, does not have ADAS, and utilizes standard light chassis (no heavy cargo surcharge).

Step 1: Calculate Expected Claim Frequency

Frequency Relativity=1.30 (Radius)×0.80 (Telematics)×1.00 (No ADAS)=1.04\text{Frequency Relativity} = 1.30 \text{ (Radius)} \times 0.80 \text{ (Telematics)} \times 1.00 \text{ (No ADAS)} = 1.04 Expected Frequency per Van=0.2000×1.04=0.2080 claims per vehicle year\text{Expected Frequency per Van} = 0.2000 \times 1.04 = 0.2080 \text{ claims per vehicle year} Total Expected Fleet Claims=50×0.2080=10.40 claims\text{Total Expected Fleet Claims} = 50 \times 0.2080 = \mathbf{10.40 \text{ claims}}

Step 2: Calculate Expected Claim Severity

Severity Relativity=1.20 (Radius)×1.00 (Standard Chassis)=1.20\text{Severity Relativity} = 1.20 \text{ (Radius)} \times 1.00 \text{ (Standard Chassis)} = 1.20 Expected Severity per Claim=$5,000×1.20=$6,000\text{Expected Severity per Claim} = \$5,000 \times 1.20 = \mathbf{\$6,000}

Step 3: Calculate Pure Premium (Loss Cost)

Pure Premium per Van=0.2080 (Frequency)×$6,000 (Severity)=$1,248\text{Pure Premium per Van} = 0.2080 \text{ (Frequency)} \times \$6,000 \text{ (Severity)} = \mathbf{\$1,248} Total Fleet Pure Premium=50×$1,248=$62,400\text{Total Fleet Pure Premium} = 50 \times \$1,248 = \mathbf{\$62,400}

(Note: To determine gross written premium, the underwriter would subsequently load this $62,400 pure premium for underwriting expenses, premium taxes, and targeted underwriting profit margin.)


Common Exam Traps in Underwriting Analytics

Caution

Trap 1: The Offset Term Misunderstanding Examination questions frequently ask about the purpose of the offset term in a Poisson frequency GLM. An offset term does not adjust for inflation, profit loads, or severity. Its sole mathematical function is to account for differing exposure sizes (e.g., car-years, payroll volume, sales) so that frequency is normalized on a per-unit basis.

Warning

Trap 2: Overfitting Symptoms If an exam question describes a machine learning model that achieves a 98% accuracy rate on training data but drops to 61% on holdout validation data, the model is suffering from overfitting (high variance). The correct remedy is to prune tree depth, apply regularization, or use cross-validation—never to add more parameters or train for more epochs.

Note

Trap 3: FCRA Adverse Action Triggers An Adverse Action Notice under the Fair Credit Reporting Act is not only required when an insurer outright cancels or denies coverage. If an applicant receives a higher price tier or is denied a preferred discount due to credit, an Adverse Action Notice is legally mandated.

Loading diagram...
Machine Learning Underwriting & Pricing Architecture
Test Your Knowledge

An actuarial team is developing a Generalized Linear Model (GLM) to price personal automobile collision coverage. In building the claim frequency model, why do the actuaries select a Poisson distribution with a log link function and include the natural logarithm of vehicle exposure years as an offset term?

A

The Poisson distribution eliminates all claim severity skewness, while the offset term forces all regression coefficients to sum to zero.

B

The log link ensures negative claim counts are possible, while the offset term inflates loss reserves for late-reported claims.

C

The Poisson distribution converts categorical rating variables into continuous distributions, while the offset term scales the baseline premium for inflation.

D

The Poisson distribution models discrete non-negative counts, the log link guarantees strictly positive expected values, and the offset term proportionally accounts for differences in policy exposure duration.

Test Your Knowledge

A data science team trains a 25-level deep decision tree to predict commercial property loss ratios. The model demonstrates a 99.2% accuracy rate on the historical training set, but when evaluated on out-of-time test data from the subsequent calendar year, accuracy plummets to 54.1%. What statistical phenomenon is occurring, and what is the appropriate remediation?

A

The model is suffering from overfitting (high variance); the actuaries should prune tree depth, implement cross-validation, or deploy an ensemble method like Random Forests or Gradient Boosting with regularization.

B

The model is suffering from underfitting (high bias); the actuaries should add more parameters, increase tree depth to 50 levels, and eliminate all regularization penalties.

C

The model is experiencing omitted variable bias; the actuaries should replace the decision tree with an ordinary least squares regression without testing on holdout data.

D

The model is exhibiting perfect generalizability; the test set must be discarded because historical training data always takes precedence over future observations.

Test Your Knowledge

Vanguard Mutual utilizes Credit-Based Insurance Scores (CBIS) to assign applicants to one of ten personal auto pricing tiers. An applicant with a clean driving record is placed in Tier 7 (a higher-premium tier) solely because their credit report reveals high credit card balance utilization and multiple recent credit inquiries. Under the Fair Credit Reporting Act (FCRA), what is the insurer legally obligated to provide to the applicant?

A

A full refund of all policy application fees, along with a certified letter from the state insurance commissioner declaring the credit score null and void.

B

An immediate waiver of all policy surcharges, because the FCRA prohibits using credit reports for applicants with clean driving records.

C

An Adverse Action Notice disclosing that credit information influenced the pricing decision, identifying the credit bureau used, and listing the primary factors that adversely impacted the score.

D

A complete copy of the insurer's proprietary predictive pricing algorithm and source code to verify model transparency.

Sections you finish are checked off in the contents.