Preparing for an AI audit means assembling six things: a model inventory, data provenance records, evaluation evidence, your human-oversight design, vendor documentation, and your incident history. None of them has to be complete. Each of them has to exist somewhere other than in a person's head.
Teams that arrive with this material in order take roughly a fortnight off the engagement clock, because evidence-chasing — not analysis — is what stretches an audit. What preparation does not do is reduce the number of findings. It usually increases it, earlier, and more cheaply.
1. A model inventory
A list of every AI system your organisation runs or relies on, with an owner's name against each one.
The list is nearly always longer than the first draft. The systems teams forget are not the models they built — those are memorable — but the AI features that arrived inside software they already had. A CRM that summarises calls, a support desk that drafts replies, a recruiting tool that ranks candidates, a security product that scores anomalies, a spreadsheet add-in that predicts. Each is an AI system your organisation is answerable for, whoever wrote it.
For each entry, five fields will carry you a long way: what it does, what decision it influences, who owns it, whether it processes personal data, and whether it is built, bought or embedded. If you can only produce this for the systems you built yourself, say so — the gap is itself a finding worth having early.
2. Data provenance
Where the training, fine-tuning and retrieval data came from, and what rights came with it.
This is the section that most often has nothing behind it, and it is the one with the longest tail if a problem surfaces after launch rather than before — a provenance issue is rarely a configuration fix. Assemble what you can: data sources and their licences, consent basis where personal data is involved, whether any of it is scraped and under what terms, what your vendors say about theirs, and what your contracts actually promise about it.
Regulators have converged on wanting organisations to be able to show this rather than assert it, which is a different standard of evidence. We cover what that means in practice in data provenance requirements, and the questions we ask about training data in the training-data questions we ask first.
3. Evaluation evidence
Whatever testing has been done, with the dates and the datasets attached.
The test is not whether the results are good. It is whether the results are traceable: what was measured, on what data, when, by whom, and against what threshold. An accuracy figure with no denominator, no date and no population attached is not evidence — it is a claim, and it will be treated as one. A bare accuracy percentage tells an auditor almost nothing.
If the honest answer is that no formal evaluation has been done, write that down. An organisation that knows it has not tested is in a considerably better position than one that believes a vendor benchmark is its own test result.
4. The human-oversight design
Not the sentence in the policy that says a human reviews the output. The actual arrangement.
Who reviews, at what stage, with what information in front of them, under what time pressure, and with what authority to overrule. This matters because "meaningful human oversight" is a load-bearing concept in the EU AI Act and in most sector rules, and because reviewers who see one decision every four seconds are not providing it. Write down the volume, the median review time, and what happens when a reviewer disagrees. If nobody has ever disagreed, that is worth knowing too.
5. Vendor documentation and contracts
Model or system cards, evaluation results, data-handling terms, subprocessor lists, and — separately — what the contract says about who is responsible for what.
These are two different documents and they frequently disagree. A vendor's marketing page will describe extensive safety testing; the contract will allocate the compliance obligation to you. The audit reads the contract. "Our vendor handles it" has not worked as a defence, and a model card is not an audit report — the two do genuinely different jobs.
6. Incident history
Every occasion the system did something it should not have, and what happened next.
Organisations are reluctant to hand this over, and it is the most valuable single item on the list. An incident log shows an auditor how the system actually fails, how failures are detected, how long detection takes, and whether anything changes afterwards. A system with no incidents recorded is either very new, very small, or not being monitored — and the third possibility is a finding.
Include near-misses and complaints. A user who contested a decision is an incident, whatever the outcome.
What to do with what you find while assembling it
Fix nothing yet, and delete nothing at all.
Preparation surfaces problems, and the instinct is to quietly correct them before the auditor arrives. Resist it for two reasons. The first is practical: a fix applied in the fortnight before an audit has not been tested, and an untested fix in a system under review creates a discrepancy between the documentation and the behaviour that costs everyone time to unpick. The second is that the record of a problem found and handled is itself good evidence. A remediated incident with a paper trail reads far better to a regulator or a counterparty than a system with no history at all.
What is worth doing is writing down what you found and what you intend to do about it. That document goes into the engagement as context and often shortens it.
The realistic timetable
A team that has never assembled this material should allow two to three weeks of part-time work — not because any single item is hard, but because five of the six require somebody else to send you something.
Start with the inventory, because everything else is organised by it, and start it before you have decided on scope: the inventory is what makes the scoping conversation short. Our own intake list, with the reasoning behind each item, is in the document list we send before we start, and what we do with it is the five scoping questions.
If you are not ready
Being unable to produce most of this is not a reason to delay. It is quite often the reason to start — an organisation that cannot answer "what AI are we running and who owns it" has a governance gap, and no amount of further preparation closes it from the inside.
The scoping conversation is free and its legitimate outcomes include "build the inventory first, then talk to us." If you would rather test the water, the free Risk Snapshot takes about a minute and produces a prioritised exposure summary from five questions.
Frequently asked questions
- What do I need before an AI audit starts?
- Six things: a model inventory, data provenance records, evaluation evidence, the human-oversight design, vendor documentation and contracts, and your incident history. None has to be perfect — the point is that it exists and that someone owns it.
- Will preparing well mean fewer findings?
- No, and it is worth being clear about that. Preparation usually surfaces findings earlier, because assembling the evidence is what reveals which claims have nothing behind them. What it changes is who finds them, when, and how expensive they are to fix.
- What if we do not have a model inventory?
- That is common and it is not a blocker. Building one is often the first deliverable of the engagement. It is worth attempting it yourself first, because the systems you forget are usually the AI features inside SaaS tools nobody thinks of as AI.
- Do we need documentation from our AI vendors?
- Yes — model or system cards, evaluation results, data-handling terms and the contractual allocation of responsibility. "Our vendor handles compliance" has not succeeded as a defence, and the audit will test what the contract actually says rather than what the sales conversation implied.
- How long does preparation take?
- For a team that has never assembled this material, allow two to three weeks of part-time work. Teams that arrive with it in order typically take about a fortnight off the engagement clock, because evidence-chasing is what stretches an audit.