Auditing AI Underwriting for Fairness: What Gets Measured

Proxy discrimination does not need the protected characteristic as an input. What a defensible test measures, why the fairness definition is a governance choice, and what carriers are asked t


By Brijesh Patel Founder & Lead Auditor
  • 8 min read
FIG. 01 Insurance · Standards explained
A dark title plate: the word "Insurance" set large in cream serif above a short gold rule, with the four outcomes a fairness test must measure — acceptance, price, terms and referral — beneath it in small gold capitals.
Insurance

An AI underwriting model does not need a protected characteristic as an input to produce discriminatory outcomes. It needs correlated variables, and it will find them — a model optimising for loss ratio will use whatever predicts loss, including the parts of a variable that are actually predicting something else.

That is the whole of the proxy-discrimination problem, and it is why "we removed the protected fields" is not a fairness control. What replaces it is measurement: outcomes by group, across the whole decision, on a schedule, against a fairness definition you have chosen deliberately.

How proxies get in

Insurance modelling has always used correlated variables, and the shift with machine learning is one of degree that becomes a difference in kind. A generalised linear model with thirty reviewed factors is something an actuary can inspect. A gradient-boosted model over several hundred features, some of them derived, some of them from external consumer data, is not — and the interactions are where the proxying happens.

Four common vectors:

  • Geography at fine resolution. Postcode-level and sub-postcode-level features carry demographic composition in nearly every market. The finer the resolution, the stronger the proxy.
  • External consumer data. Purchased attributes about behaviour, lifestyle, digital activity and household composition are the vector most directly targeted by state legislation, and the one carriers have least visibility into. A vendor attribute is a variable whose construction you did not see.
  • Behavioural and channel signals. Application channel, device, time of day, session behaviour and how a quote was reached all correlate with demography.
  • Interactions. Two individually innocuous features can jointly reconstruct a protected characteristic, and this is invisible to any review that looks at features one at a time.

What a defensible test measures

The most common inadequate test measures one thing at one place: acceptance rate at the decline threshold. In most books of business, the exposure is not there.

A test that will stand up measures outcomes across the whole decision:

  • Acceptance — the decline rate by group, which is the part everyone measures.
  • Price. Two applicants who are both accepted at materially different premiums have been treated differently, and pricing dispersion is where a great deal of unexamined exposure sits.
  • Terms. Exclusions, excesses, limits and conditions are part of the outcome.
  • Referral and manual-review rates. A group referred to manual underwriting at three times the base rate is being treated differently even if the eventual acceptance rates converge — the process burden is itself the differential.
  • Score distribution, not just the pass rate. A shifted distribution tells you the exposure will move the moment the threshold does.

And it states its method before its results: which populations, constructed how, with what sample sizes, on what date, against what threshold for materiality. A test whose thresholds are set after the numbers are known is not a test. What a defensible fairness evaluation measures goes through the construction in detail.

The definition is a governance choice

This is the part most often skipped, and it is unavoidable.

The standard statistical fairness definitions — equal acceptance rates across groups, equal error rates, equal calibration — are mutually incompatible in essentially every real dataset where base rates differ. It is not a matter of trying harder or modelling better; satisfying one generally means failing another. A carrier that has not chosen a definition has chosen one implicitly, by whatever its model optimised for.

The defensible position is to choose deliberately, record the reasoning, and be able to explain it to a regulator, a reinsurer or a board. That record is worth more than a good number, because a good number under an unstated definition cannot be interrogated — and an examiner asking "why this measure" is asking a question that has no good improvised answer.

Claims is the less-examined half

Underwriting attracts the fairness attention; claims automation frequently does not, and it carries the same class of exposure with less monitoring around it.

A claims-triage model affects how quickly a claim is paid. A fraud-scoring model affects whether a claimant is investigated, how long that takes, and what it feels like to be on the receiving end. A settlement-recommendation model affects the amount. Differential outcomes in any of those are differential treatment, and the standard defence — that the model only recommends and a human decides — is only as good as the human review behind it.

The questions worth asking of your own claims models: what is the reviewer's caseload per hour, what information do they see, how often do they depart from the recommendation, and have the departure rates ever been analysed by claimant segment. If nobody has ever departed from the recommendation, the human step is nominal.

What carriers are asked to document

The NAIC's model bulletin on the use of AI systems by insurers sets an expectation that a carrier maintains a written AI systems programme: governance and accountability, risk management and internal controls across the AI lifecycle, and oversight of third-party AI and data. Adoption varies by state and continues to change, so the current position has to be checked in each state you write in rather than taken from any summary, including this one.

Colorado's SB21-169 goes further in a specific direction, restricting insurers' use of external consumer data and information sources and the algorithms built on them where they result in unfair discrimination, with an associated testing and reporting regime. It is the instrument most often raised with us because it targets exactly the vector carriers have least visibility into.

Where a carrier writes EU business, the EU AI Act adds its own classification question, answered per system against the primary text rather than by sector assumption. And the NIST AI Risk Management Framework, though voluntary, is increasingly the shape reinsurers and enterprise partners use when they ask how AI risk is managed.

Third-party models

A large share of insurance AI is vendor-supplied, and the obligation does not transfer with the model. What a carrier needs from a vendor is specific: the variable list, including derived features; the population the model was developed and validated on; fairness testing results broken out by segment; whether external consumer data is used and from which sources; the change-notification terms; and what the contract actually allocates.

That last item is the one to read first. Vendor marketing describes extensive testing; vendor contracts generally allocate the regulatory obligation to the carrier. The audit reads the contract.

What to have in place

Four things carry most of the load: a written fairness definition with the reasoning recorded; a testing protocol covering acceptance, price, terms and referral, on a schedule rather than at launch; a variable review that examines interactions and derived features rather than a flat list; and monitoring with pre-set thresholds and a named person obliged to act when one is crossed.

Findings from this work are rarely quick fixes — a proxy embedded in a model's feature interactions is not a configuration change — which is why the roadmap matters as much as the report. Turning findings into a fix roadmap that sticks covers that half.

How we approach the domain is set out on the insurance page, and the standard our own reports are measured against is published in full as the verification methodology.

Frequently asked questions

What is proxy discrimination in insurance AI?
A model producing materially different outcomes for a protected class using variables that are not themselves protected characteristics but correlate with them. It does not require intent, and removing the protected variable does not prevent it — a model will reconstruct the signal from whatever correlates with it.
What does a defensible fairness test measure?
Outcomes by group across the whole decision, not just the accept/decline boundary: acceptance, pricing, terms, referral rates and the score distribution. It states its populations, its method, its thresholds and its date before it reports results, and it is repeated on a schedule rather than once at launch.
Which fairness definition should we use?
That is a governance decision, not a technical one. The common statistical definitions are mutually incompatible in most real datasets — you cannot satisfy all of them at once — so the defensible position is to choose one, write down why, and be able to explain the choice. A carrier that has not chosen has chosen implicitly.
What do insurance regulators ask carriers to document?
The NAIC model bulletin on the use of AI systems by insurers sets an expectation of a written AI systems programme covering governance, risk management and controls, and third-party AI oversight. Adoption varies by state and continues to change, so check the current position in each state you write in rather than relying on any summary.
Does this apply to claims automation as well as underwriting?
Yes, and claims is often the less-examined half. A claims-triage or fraud-scoring model affects payout speed, investigation burden and settlement outcomes, and differential treatment there is the same kind of exposure as in underwriting — with less monitoring around it, in our experience.

Where your AI stands

Wondering where your AI stands?