OSFI GUIDELINE E-23 · MODEL RISK MANAGEMENT · EFFECTIVE 1 MAY 2027

Model risk management, built for AI and ML.

E-23 sets OSFI's expectations for model risk management across every federally regulated financial institution, over the whole model lifecycle, from 1 May 2027. The model definition expressly reaches AI and machine learning, insurers are in scope for the first time, and non-financial-risk models come with them. We give you the inventory, the risk ratings, the independent review and the gap plan.


A bound model file on a dark desk beside reading glasses, a fountain pen and a card reading evidence over promises Illustrative materials
A supervisor asks for the file, not for the intention
Model inventory Risk rating Independent review Five lifecycle stages AI and ML in scope
The guideline

What OSFI Guideline E-23 is

A principles-based supervisory expectation, technology-neutral and applied on a risk basis — which means proportionate, not optional.

Guideline E-23 was published by the Office of the Superintendent of Financial Institutions on 11 September 2025 and takes effect on 1 May 2027. It sets OSFI’s expectations for how a federally regulated financial institution identifies, measures and controls the risk that comes from using models — across the whole lifecycle, not only at build and validation.

The definition is the part that catches people. A model is an application of theoretical, empirical, judgmental assumptions or statistical techniques — including AI and machine learning methods — that processes input data to generate results, in three components: data input, processing, results. It does not turn on what your team calls the thing, or on the technology it runs in.

Two expansions matter. Insurers are in scope — life, fraternal, and property and casualty — alongside banks, foreign bank branches and trust and loan companies. And the reach now includes non-financial-risk models: climate, cyber, technology and digital innovation, not just the models that produce a capital number.

What E-23 asks for is an enterprise-wide framework aligned to your risk appetite, a model inventory covering everything of non-negligible risk, named roles from owner to approver, a risk rating that actually drives how much governance each model gets, and independent assessment of conceptual soundness and performance — applied across five lifecycle components.

Effective 1 May 2027

Published 11 September 2025. Discovery is the phase that overruns, so the runway is shorter than it looks.

Applies On a risk basis

Proportionate to size, strategy, risk profile, nature, scope and complexity — and the proportionality argument is itself evidence.

Method Technology-neutral

The same rigour applies to a logistic regression and to a large model nobody can fully explain. Neither is excluded, and neither is excused.

Who needs to comply

  • Banks and foreign bank branches

    Domestic banks and the Canadian branches of foreign banks. Foreign entities comply to the extent consistent with their applicable requirements and legal obligations under Guideline E-4.

  • Trust and loan companies

    In scope on the same terms as banks, with the same proportionality allowance.

  • Life insurance and fraternal companies

    New to E-23 with this version. Pricing, reserving, capital and underwriting models all sit inside the definition.

  • Property and casualty companies

    Also new. Rating, reserving, catastrophe and claims models are in scope wherever their output is non-negligible.

  • Anyone running non-financial-risk models

    The definition reaches climate, cyber, technology and digital-innovation models, not only the models that produce a capital number.

  • Pension plans — check before assuming

    OSFI consulted on extending E-23 to federally regulated private pension plans. The guideline as published addresses FRFIs; if you administer a plan, confirm your position rather than reading the consultation as settled.

Risk rating

Tier every model, then calibrate the rigour

E-23 expects the rating to come from quantitative factors — portfolio size, financial impact — and qualitative ones: business purpose, complexity, data reliability, customer impact, regulatory risk. What the rating then buys you is a proportionate amount of control.

Tier 1

Highest materiality

Models whose failure would move capital, solvency, pricing at scale or a regulatory submission.

  • Full independent validation before use
  • Annual revalidation at minimum
  • Board-visible reporting
  • Tight change control and re-approval
Tier 2

Moderate materiality

Meaningful business impact, but contained — a single portfolio, product line or process.

  • Independent review, scoped to risk
  • Periodic revalidation on a defined cycle
  • Monitoring with escalation thresholds
  • Documented approval to deploy
Tier 3

Lower materiality

Limited impact, easily overridden, or used only to inform a decision a person still makes.

  • Proportionate review, not full validation
  • Inventory entry and owner named
  • Light monitoring against expectations
  • Re-rated if the use changes

Three tiers is a pattern, not a requirement. E-23 requires a risk rating built from the factors above, and expects it to drive the intensity of governance proportionally. It does not prescribe how many bands you use or where the boundaries sit. Most institutions land on three; if you do, be ready to say why your boundaries are where they are — that answer is the control, not the label.

Expectations

What E-23 expects you to do

Eight principles-based expectations, from an enterprise-wide framework down to the AI and ML controls. The chips are themes, not citations — E-23 does not number its expectations.

Governance

An enterprise-wide MRM framework

Risk-based policies to identify, assess, manage and monitor model risk, proportionate to size, complexity and interconnectedness — one framework across the institution, not a policy per business line that agrees with no other.

Scope

A model definition that reaches AI and ML

Theoretical, empirical, judgmental or statistical — including AI and machine learning methods — processing input data to generate results. If it produces an output someone relies on, start from the assumption that it is a model.

Lifecycle

Governance at every lifecycle stage

Design, review, deployment, monitoring and decommission each carry expectations. Programmes that cover development and validation and stop there are covering two of five.

Rating

Risk rating that drives the rigour

An inherent-risk rating from quantitative factors — portfolio size, financial impact — and qualitative ones: business purpose, complexity, data reliability, customer impact. Refreshed on trigger events, not once a year.

Validation

Independent assessment

Validation independent of development, confirming sound specification and fitness for purpose. Independence is of judgement, not merely of reporting line.

Accountability

Named roles across the model

Model owner, developer, independent reviewer, approver and user, each defined. Senior management holds enterprise accountability; the board hears about model risk.

Inventory

A firm-wide model inventory

Every model of non-negligible risk with id, owner, version, rating, data sources, dependencies, approved uses, limitations, review dates and decommission status — and vendor models on it too.

Monitoring

Ongoing monitoring and AI/ML controls

Performance watched against expectations after deployment, plus transparency and explainability, controls for black-box or autonomous models, and bias, privacy and drift monitoring under multi-disciplinary governance.

The lifecycle

The five lifecycle components

E-23 governs the model from idea to retirement. Controls and documentation are expected at every stage — and the last one is where nearly every existing programme stops.

01

Design

Purpose, data, assumptions and method chosen and documented, with the limitations written down while they are still obvious to the people who chose them.

  • Business purpose and intended use
  • Data lineage, quality and fitness
  • Method selection and rejected alternatives
  • Known limitations and out-of-scope uses
02

Review

Independent assessment of conceptual soundness and performance before the model carries weight, scoped to the model’s risk rating.

  • Conceptual soundness challenged, not confirmed
  • Testing against the intended use
  • Findings tracked with owners and dates
  • A recorded approval decision
03

Deployment

Controlled release into the environment it will actually run in, with the difference between the tested model and the deployed model closed rather than assumed away.

  • Approval before use, on the record
  • Implementation testing in the live path
  • Change control and version pinning
  • User guidance and training
04

Monitoring

Performance watched against expectations, with thresholds that fire and a route from a breached threshold to a decision.

  • Metrics defined before go-live
  • Drift, stability and data-quality checks
  • Escalation thresholds with named owners
  • Periodic re-rating as use changes
05

Decommission

The stage almost every existing programme is missing. A model that has been switched off still has downstream consumers, retained outputs and an inventory entry — retiring it is a controlled act, not an absence of activity.

  • Retirement decision and rationale recorded
  • Downstream dependencies identified and told
  • Data and output retention resolved
  • Inventory status closed, not deleted
Consequences

How compliance is enforced

E-23 is a guideline, not a penalty schedule. OSFI is a prudential supervisor, and the consequence of a weak model risk programme arrives as supervisory pressure that escalates.

Stage 1

Supervisory findings

Issues raised through ongoing supervision, with remediation expected on OSFI’s timetable rather than yours. This is where most institutions meet E-23 for the first time.

Stage 2

Heightened scrutiny

Weaknesses in model risk feed the assessment of governance and risk management overall — and a rating that moves the wrong way brings more supervisory attention across the board.

Stage 3

Intervention and capital

Under OSFI’s staged intervention framework, persistent weakness can mean restrictions on activity or a capital consequence. That is the end of the escalation, not its start.

No fine appears on this page because there is none to quote. OSFI supervises rather than penalises, and the pressure it applies — findings on its timetable, a weaker view of your governance, and ultimately the staged intervention framework — is harder to discharge than a payment would be.

Expectation & coverage

E-23 expectations, mapped to what we build

Twelve expectations, and the artefact that discharges each one. It helps you meet the guideline — it does not replace your own judgement, and every claim here can be checked against the expectation beside it.

E-23 expectation What we build
Enterprise MRM framework Governance A framework document aligned to your risk appetite statement, with the scope boundary written down — including what you have decided is not a model, and why.
Model definition and scoping test Scope A definition your teams can actually apply, with worked examples on the borderline cases: rules engines, spreadsheets, vendor scores, LLM-assisted steps.
Model inventory Inventory An enterprise register with owner, rating, purpose, status, dependencies and vendor provenance — and a route by which a new model actually gets onto it.
Risk rating methodology Rating A rating scheme built from both the quantitative and qualitative factors, applied across the inventory, with the borderline calls documented.
Roles and accountabilities Accountability Owner, developer, reviewer, approver and user defined and mapped to named people, plus the board and senior-management reporting line.
Lifecycle standards Lifecycle A standard per stage — design, review, deployment, monitoring, decommission — with the evidence each stage has to leave behind.
Independent review function Validation A review methodology scaled by rating, a findings register with owners and dates, and the independence argument written rather than asserted.
Monitoring and thresholds Monitoring Metrics per model with thresholds, owners and an escalation route — and for AI and ML, drift and data-quality checks that run without being asked.
Decommission procedure Lifecycle The retirement runbook: dependency check, retention decision, notification list and the inventory status that closes rather than disappears.
Vendor and third-party models Scope The documentation you need from a supplier before a bought model can be rated, reviewed and monitored like one of your own.
Board and management reporting Governance A reporting pack that gives the board model risk in a form it can act on, at a cadence that survives a quiet quarter.
Gap plan to 1 May 2027 Readiness A findings list against every expectation on this page, sequenced, owned and dated against the effective date.

See the mapping against your own model estate

Gaps

Where most model risk programmes fall short

Teams that have been running model risk for years still meet E-23 with these eight. The first two are scope failures, and they are the ones that make the rest invisible.

The inventory is a spreadsheet

Maintained by one person, updated when someone remembers, and unreconciled against production. An inventory that cannot be trusted is worse than none, because it makes the gaps invisible.

Half the models are not called models

Rules engines, pricing spreadsheets, vendor risk scores and an LLM step in a workflow. E-23’s definition turns on what the thing does, not on what the team calls it.

Validation is a document review

Reading the developer’s documentation and agreeing with it is not an assessment of conceptual soundness. Independence of reporting line without independence of judgement buys nothing.

Monitoring is annual, in a slide

Performance reviewed once a year in a governance pack is not monitoring. Without a threshold that fires, nobody finds out until a business outcome tells them.

Nothing is ever decommissioned

Models are switched off without a retirement decision, downstream consumers keep reading stale outputs, and the inventory carries entries nobody can account for.

The rating does not change anything

Every model rated, and every model then governed identically. If Tier 3 gets the same treatment as Tier 1, the rating is a label rather than a control.

Vendor models are taken on trust

A bought model is your model risk. Without documentation on data, method and limitations you cannot rate it, review it or monitor it — and the supervisor asks you, not the supplier.

The board hears a colour, not a risk

Reporting that reduces the estate to a heat map tells the board nothing it can act on. E-23 expects model risk to reach the board in a form that supports a decision.

How we help

How iDharma supports OSFI E-23

Six workstreams — discovery, rating, independent review, monitoring, lifecycle standards and the gap plan. Each one is a change in how the institution operates, not a document that lands and then ages.

Model discovery and inventory

Found from the systems inwards rather than the questionnaire outwards — including the spreadsheets, the vendor scores and the model nobody calls a model.

Risk rating and tiering

A rating scheme applied across the estate, so the intensity of your controls can be defended against the factors E-23 actually names.

Independent review

Conceptual soundness and performance assessed by people with no stake in the answer — which is the whole point of the word independent.

Monitoring design

Metrics, thresholds and escalation defined before go-live, with drift and data-quality checks for the AI and ML models that need them.

Lifecycle standards and decommission

All five stages written as standards with an evidence requirement each — including the retirement runbook nearly every programme is missing.

Readiness gap plan

Findings against every expectation, sequenced against 1 May 2027, with owners and dates rather than a heat map.

Every decision above leaves a record. A rating, a review conclusion, a breached threshold and a retirement are each dated and attributable by the time we are finished — which is the reproducible trail E-23’s documentation expectations are actually asking for.

AI and ML

E-23 puts AI and ML directly in scope

The model definition names them. The guideline stays technology-neutral, which means an opaque method is neither excluded nor excused — three consequences follow.

Transparency and explainability

The guideline is technology-neutral, which cuts both ways: an opaque method is not excluded, and it is not excused either. The harder the model is to explain, the more the rest of the framework has to carry.

Drift, bias and data quality

A model that was sound at review can stop being sound without anything changing in the code. Monitoring has to watch the inputs and the population, not only the outputs.

Multi-disciplinary governance

Reviewing an AI model needs data science alongside risk, legal and the business. A validation function staffed only by traditional quants will pass models it cannot actually assess.

The target state

What E-23 ready looks like

Four tests, phrased the way a supervisor asks them. Deliberately not four percentages — a coverage score in this position is a claim about a product, and none of these four is ours to score.

Inventory

Complete and reconciled

Every model of non-negligible risk on one register, reconciled to production, with an owner who knows they own it.

Rating

Applied and consequential

A rating on every entry, built from the named factors, visibly driving how much review and monitoring each one gets.

Review

Independent and current

Tier 1 validated and in date, findings tracked to closure, and the independence argument written down.

Monitoring

Running and escalating

Thresholds defined per model, breaches routed to a named owner, and at least one instance where that actually happened.

Questions

Frequently asked questions

Scope, the model definition, tiering and what actually happens if you are not ready.

1 Scope and timing
Who does Guideline E-23 apply to?

Federally regulated financial institutions: banks, foreign bank branches, trust and loan companies, life insurance and fraternal companies, and property and casualty companies. Foreign entities comply to the extent it is consistent with their applicable requirements and legal obligations under Guideline E-4.

When does it take effect?

It was published on 11 September 2025 and takes effect on 1 May 2027. That is not a long runway for an institution that has to build an inventory first — discovery is the phase that overruns.

What changed from the earlier version?

The reach. The previous guideline was aimed at deposit-taking institutions and models used for capital; this one covers all FRFIs including insurers, expressly names AI and machine learning methods inside the model definition, and extends to non-financial-risk models such as climate, cyber and technology.

Are pension plans in scope?

OSFI consulted on extending E-23 to federally regulated private pension plans, and the guideline as published addresses FRFIs. If you administer a plan, confirm your position with your own advisers rather than treating the consultation as settled.

We are small. Does all of this apply?

It applies on a risk basis, proportionate to your size, strategy, risk profile, nature, scope and complexity. That means a smaller programme, not an absent one — and the proportionality argument itself is something you should be able to show.

2 Models, ratings and validation
What counts as a model?

An application of theoretical, empirical, judgmental assumptions or statistical techniques — including AI and machine learning methods — that processes input data to generate results. Three components: data input, processing, results. If something produces an output a decision relies on, start from the assumption it is in scope.

Is a spreadsheet a model?

It can be. The definition turns on what the thing does, not on the technology it runs in. A spreadsheet applying judgmental assumptions to produce an estimate someone acts on is a model; a spreadsheet that adds up a column is not.

Does E-23 require three tiers?

No. It requires a risk rating driven by quantitative and qualitative factors, and expects that rating to drive the intensity of governance. Three tiers is the pattern most institutions land on, not a requirement — and if you use it, you should be able to say why the boundaries sit where they do.

How independent does validation have to be?

Independent enough to assess conceptual soundness and performance without a stake in the answer. A separate reporting line helps and is not sufficient on its own; the test is whether the reviewer can and does disagree with the developer.

Do vendor models count?

Yes, and they are yours to govern. You need enough from the supplier on data, method and limitations to rate the model, review it and monitor it. Where the vendor will not provide it, that gap is a finding against you.

3 Enforcement and the engagement
What happens if we are not ready?

OSFI does not fine. The consequence is supervisory and it escalates: findings raised through ongoing supervision, then a weaker assessment of governance and risk management overall, and in a persistent case the staged intervention framework, which can reach restrictions on activity or a capital consequence.

Does E-23 sit alongside other frameworks?

It maps cleanly onto NIST AI RMF and onto SR 11-7 and SS1/23 if you already run those for other jurisdictions. Most of the inventory, rating and validation work is common; what differs is the lifecycle framing and the supervisory audience.

Where do most readiness programmes lose time?

Discovery. Institutions consistently underestimate how many models exist outside the places they already look, and every downstream activity — rating, review, monitoring — is blocked until the inventory is trustworthy.

What does an iDharma E-23 readiness review cost and how long does it take?

It is scoped before you are charged. The variables are the size of the estate, how much of the inventory already exists, and whether AI and ML models are in the population; we tell you the shape of all three after a short scoping call.

Further reading

Read it at source

The dates, the model definition and the five lifecycle components on this page came off OSFI’s own guideline. Where a scoping decision turns on the wording, go to it.

Ready when you are

Get E-23 ready before 1 May 2027

A readiness review gives you the model inventory, the risk ratings, the lifecycle standards, the independent-review methodology and a sequenced gap plan against the effective date — scoped before you are charged.

This page is guidance on how we scope an E-23 readiness review, not legal or supervisory advice. E-23 is principles-based, so more of it is arguable than a rules-based text would be; where a call is genuinely open we say so in writing rather than pick the convenient answer.