AI in institutional finance raises a different set of questions from AI in consumer credit. Trading, portfolio-allocation, AML-monitoring and financial-reporting systems are rarely making a decision about an identifiable person, so the fair-lending framing that dominates lending work does not transfer. What replaces it is a harder question about evidence: can you reconstruct why the system did what it did, on a specific day, on a specific input?
An audit here is largely an examination of whether that reconstruction is possible.
What comes into scope
Four families of system, with different centres of gravity.
Trading and execution models raise reconstruction and control questions: can a specific decision be explained after the fact, what constrains the model's behaviour in conditions it has not seen, and what stops it. The control layer around the model is usually more consequential than the model.
Portfolio and allocation models raise suitability and drift questions. A model tuned on one regime and running in another is the standard failure, and it is invisible from inside the model's own metrics.
AML and surveillance systems raise a distinctive problem: they fail silently. A monitoring system that stops catching a pattern produces no alert saying so, and the only way to know is to test what it misses rather than to measure what it flags. Tuning history and the review trail behind threshold changes matter more here than almost anywhere else.
Financial-reporting and forecasting AI raises lineage and reproducibility questions. If a number reached a report through a model, the audit trail has to reach back through the model to the data, and generative components in that chain — a model summarising, extracting or classifying — are frequently undocumented because nobody classified them as a model.
That last point is worth generalising. The AI in a financial institution that is least likely to be in the inventory is the AI that arrived inside software already licensed: summarisation in a research platform, extraction in a document pipeline, classification in a workflow tool. Each is a model in a decision chain, and each is in scope of the same expectations as a model somebody built.
The four evidence gaps that surface most often
1. Validation independence, satisfied nominally
The supervisory expectation — SR 11-7 is the reference point most examiners and institutional counterparties use — is that a model used in decision-making is validated by someone independent of the people who developed or selected it. The expectation is technology-neutral; an AI model is a model.
What we find is that independence is frequently satisfied by role title rather than by reporting line. A validation function that reports, eventually, to the person who owns the model's performance is not independent of it, and this is the question a counterparty is actually asking when it asks who validated. It is also the question a self-assessment can never answer well, which is the distinction between independence and self-assessment.
2. Explanations that are not faithful
Institutions under pressure to explain model outputs increasingly generate explanations with attribution methods bolted onto the model after the fact. Very few test whether those explanations are faithful — whether the features named as drivers actually drove the output, or merely correlate with it.
An unfaithful explanation is worse than no explanation, because it is relied upon. The audit question is not "can you explain it" but "what evidence do you have that the explanation is correct." Regulators and engineers mean different things by explainability, and the gap is where this failure lives.
3. Data lineage that stops at the warehouse
Lineage documentation in finance is usually strong up to the point the data enters the model and absent thereafter. What was actually in the training window, which vendor feeds were included and under what licence, what was excluded and why, how the target was defined, and whether any of it leaked information unavailable at decision time — these are the questions that determine whether a validation result means anything.
Target leakage deserves a specific mention because it produces excellent backtests and poor live performance, and it is very hard to find from outside the data. What regulators expect you to be able to prove about provenance covers the documentation side.
4. Monitoring nobody is obliged to act on
Most institutions monitor. Fewer have thresholds set in advance, and fewer still have a named person obliged to do something when one is crossed. A dashboard without an escalation duty is telemetry — it records the failure, it does not catch it.
The audit test is simple and rarely passed cleanly: name the last time a monitoring threshold was breached, and show what happened next.
Which frameworks apply
NIST AI RMF has become the common vocabulary in US financial-services questionnaires. It is voluntary, and it is useful precisely because it is not a checklist — it asks whether a process exists, whether it produces evidence, and whether anyone acts on the evidence. What the framework means in practice for this sector is a separate read, as are its four functions in plain English.
ISO/IEC 42001 is the management-system standard, and it is what an institutional counterparty is reaching for when it asks whether your AI governance will still exist next year. It describes machinery, not models.
The EU AI Act classifies by use case rather than by sector. A good deal of institutional trading and portfolio AI sits outside its high-risk categories while creditworthiness assessment sits inside — which means the classification has to be done per system, against the primary text, rather than assumed from the sector. Its obligations phase in on a timetable that has been subject to amendment proposals; verify the current schedule at source.
Sector supervision sits on top of all three and is the part no article should generalise about. Which specific rules bind a specific entity is a question for counsel, and an audit's job is to produce the evidence those rules ask for, not to opine on their application.
Where consumer credit splits off
If the system in question assesses creditworthiness or sets terms for an identifiable individual, the centre of gravity moves: adverse-action reasons and disparate-impact testing become the dominant questions, and a different body of rules applies. That work is covered in AI model risk management for lenders.
Many institutions run both kinds of system and treat them as one governance problem. They are not — the evidence a counterparty asks for is materially different, and an inventory that does not distinguish them tends to produce the wrong evidence for both.
How we approach the domain is set out on the finance page, and the standard our own reports are measured against is published as the verification methodology.
Frequently asked questions
- What does an AI audit cover in financial services?
- The model inventory and who owns each entry; the independence of whoever validated each model; whether an individual output can be explained and reconstructed; the data lineage feeding the system; monitoring and drift controls; and the human-override design. The emphasis differs by system, but that set is common to trading, portfolio, AML and reporting AI.
- Is a trading model in scope of the EU AI Act?
- It depends on what it does and who it affects. The Act's high-risk categories are defined by use case rather than by sector, and much institutional trading and portfolio AI falls outside them — while creditworthiness assessment and certain employment and access uses fall inside. This is a question to answer against the primary text for each specific system, not by sector assumption.
- How is an AI audit different from model validation?
- Model validation asks whether a model is fit for its stated purpose. An AI audit asks that too, and then asks a set of questions validation traditionally does not: data provenance and rights, adversarial robustness, whether the human oversight is real, whether the explanations given are faithful, and whether the governance around the model would catch a failure.
- Does AI in AML monitoring need special treatment?
- It attracts a particular kind of scrutiny, because a monitoring system that fails quietly produces no signal that it has failed. The questions that matter are how the system was tuned, what it was tested against, whether anyone has measured what it misses rather than only what it flags, and how a change to thresholds is reviewed and recorded.
- Who should validate an AI model in a financial institution?
- Someone independent of the people who built or selected it, with a reporting line that does not run through the model owner. That independence is the substance of the requirement rather than a formality, and it is the part most often satisfied on paper only.