9.1 Communicating Methodology & Modeling Decisions

Key Takeaways

  • Task C-2 requires communicating project technical information including methodologies, modeling decisions, and interpretation of output — three separate obligations, not one.
  • A technical reader needs the target variable, distribution, link, offset or weight, and data basis stated explicitly before any result can be evaluated.
  • Describing the iterative build — what was tried, which diagnostic prompted the change, what happened — is a distinct rubric criterion from describing the final model.
  • Model output must be interpreted, not merely displayed: a coefficient table without a sentence saying what the relativities mean does not satisfy C-2.
  • Limitations and residual uncertainty belong in the technical narrative; omitting them reads as a claim that none exist.
Last updated: September 2026

Three Obligations, Not One

Task C-2 in the PCPA Content Outline reads: "Communicate project technical information, including details on methodologies, modeling decisions, and interpretation of output." Three things are listed, and the project rubric turns them into four separate criteria:

  • Describes what method(s) they used and why.
  • Describes which model they selected and why.
  • Explains what they did to iteratively check and build their model, including which diagnostics they used and why.
  • Generates and provides technical output, such as model coefficients.

A report that only presents the final model satisfies one of these. This is where candidates most often lose marks that had nothing to do with modelling skill.

The Methodology Facts a Technical Reader Needs

A reviewing actuary cannot evaluate a result without knowing how it was produced. State these explicitly and early:

FactExample wording
Target variable"Claim frequency, measured as claim count per earned car-year."
Unit of analysis"One record per policy term."
Data basis"Policy years 2021-2024, case-incurred losses including ALAE, evaluated as of 30 June 2026."
Distribution and link"Poisson GLM with a log link."
Offset or weight"Log of earned exposure entered as an offset."
Exclusions"Records with zero or negative earned exposure removed (n = 412)."
Validation scheme"70/30 random split, stratified on policy year; 5-fold cross-validation on the training set."

Seven short statements. In a 1,250-word report they cost perhaps 90 words, and without them every number that follows is unverifiable.

Narrating the Iterative Build

The rubric wants the path, not just the destination. The pattern that works is a short causal chain repeated for each significant decision:

Observation → diagnostic → action → result

For example:

"The one-way plot of driver age showed a U-shaped relationship with frequency, so a single linear term was replaced with a five-band categorical. Deviance fell by 214 on 4 degrees of freedom and AIC improved from 48,102 to 47,896."

That sentence carries the diagnostic used, why it was used, the modelling change, and the evidence that the change helped. Three or four such sentences give a grader the whole build.

Things worth narrating this way:

  • A distribution or link changed after residual review.
  • A predictor banded or transformed after a partial-residual or one-way plot.
  • A variable dropped after a VIF or correlation finding.
  • An interaction added after a two-way exhibit.
  • A candidate model rejected because out-of-sample lift did not hold.

Dead ends are worth reporting too. "A Gamma severity model with an inverse link was tried; residual diagnostics were materially worse than the log link, so the log link was retained" demonstrates exactly the judgement being assessed.

Interpreting Output Rather Than Displaying It

A coefficient table is technical output. It is not yet interpretation. Every exhibit needs at least one sentence saying what it means.

Raw displayInterpreted
"Territory 3 coefficient = 0.336.""Territory 3 shows a frequency relativity of 1.40, meaning risks there generate about 40% more claims per car-year than the base territory, holding other rating variables constant."
"AIC: Model A 47,896; Model B 47,903.""AIC marginally favours Model A, but the 7-point gap is small relative to the 214-point improvement from banding age, so Model A was selected on the strength of its out-of-sample lift rather than AIC alone."
"Dispersion parameter = 2.7.""A dispersion estimate of 2.7 indicates substantial overdispersion relative to the Poisson assumption, so standard errors from the Poisson fit understate uncertainty and were adjusted."

The second column is what the rubric means by interpretation of output. Note that each one also says what the reader should do with the fact.

Two Phrases to Use Precisely

"Holding other variables constant." This is what distinguishes a GLM relativity from a one-way average, and it is the single most useful phrase in a technical narrative. It also explains to a reviewer why the model's territory relativity differs from the one-way indication.

"Statistically significant" is not "material." With 400,000 records almost everything is significant. A predictor with a p-value below 0.001 and a relativity range of 0.99 to 1.01 does not belong in a rating plan. Say which test you applied and why the effect matters in dollars, not only whether it passed a threshold.

Stating Limitations

A technical communication that reports no limitations implies there are none. Every model has them, and naming them is a mark of competence rather than weakness:

  • Thin data in specific segments, and what was done about it.
  • Data defects that could not be resolved, and their likely direction of effect.
  • The range over which the model should be applied, and where it should not be extrapolated.
  • Assumptions that would need revisiting if conditions change.

[!WARNING] Do not describe a GLM as "proving" a causal relationship. A GLM estimates association conditional on the other predictors in the model. The correct technical language is that the model estimates the expected value of the target given the predictors, and that observed relativities may reflect unmodelled characteristics correlated with the rating variable.

Test Your Knowledge

Which report sentence best satisfies the rubric criterion that the candidate explain the iterative model build, including which diagnostics were used and why?

A
B
C
D
Test Your Knowledge

A report presents a table of fitted coefficients with no accompanying commentary. Which C-2 obligation is unmet?

A
B
C
D
Test Your Knowledge

With 400,000 records, a predictor has a p-value below 0.001 but its fitted relativities span only 0.99 to 1.01. What is the appropriate technical communication?

A
B
C
D