What Is an AI Audit? Scope, Standards, and What You Get

An independent review of what your AI actually does, measured against a named standard — not a certificate, and not a review of what the documentation says it does.


By Brijesh Patel Founder & Lead Auditor
  • 7 min read
FIG. 01 Methodology notes
A magnifying glass resting on a printed audit report, its lens enlarging the SCOPE and METHOD headings, with a FINDINGS table and bar charts alongside it on a dark desk under low warm light.

An AI audit is an independent, evidence-based review of an AI system against a named standard. It tests what the system actually does in operation, rather than what its documentation says it does. What it produces is a written record: the scope that was reviewed, the evidence that was examined, the findings ranked by severity, and a path to fixing them.

It is not a certificate, and it is not a guarantee of compliance. Those two sentences do most of the work of this article, and the rest of it explains why.

What an AI audit actually examines

Most of the confusion about AI audits comes from the word "audit" doing two different jobs. A financial audit tests assertions against records. An AI audit tests assertions against behaviour — which means the system has to be run, not just read about.

A review that deserves the name looks at four things.

The system, as deployed

Not the model in the paper, and not the model on the vendor's benchmark. The version that is running, with the prompts, thresholds, filters, retrieval sources, fallbacks and human review steps that are actually wired around it. An enormous share of real AI risk lives in that wrapper rather than in the model weights — a well-behaved model behind a retrieval layer pointed at a stale document store produces confidently wrong answers all day.

The evidence behind the claims

Every AI system arrives with claims attached: an accuracy figure, a bias test, a security review, a statement that a vendor "handles compliance." An audit asks what each claim was measured on, by whom, when, and whether the measurement would still hold on today's traffic. This is usually where the first findings come from, and it is the reason a single accuracy number tells you almost nothing on its own.

The decisions the system influences

Scope follows consequence. A model that drafts internal meeting notes and a model that scores loan applications are both "AI," and no sensible review gives them the same treatment. What matters is whether an output changes an outcome for a person — credit, care, employment, access, price — and how much human judgement genuinely sits between the model and that outcome. "A human reviews it" is a claim like any other, and an audit tests whether the reviewer has the time, the information and the authority to disagree.

The record you would have to produce

Regulators, partner banks, enterprise buyers and courts all ask a version of the same question: show me what you knew, when you knew it, and what you did. That means a model inventory, data provenance, evaluation results, incident history, and the decision trail behind the risk classification. Organisations are very often surprised by how much of this exists in someone's head or a Slack thread rather than in a document.

The standards an audit is measured against

An audit without a named standard is an opinion. The standard is what makes a finding arguable — it gives you something to point at and, just as importantly, something to point back at.

The EU AI Act

Regulation (EU) 2024/1689 is law, not guidance, and it reaches any AI system whose output is used in the EU regardless of where the organisation running it sits. It works by risk tier: a small set of prohibited practices, a defined "high-risk" category carrying substantial obligations around risk management, data governance, documentation, logging, human oversight and accuracy, and lighter transparency duties elsewhere. Its obligations phase in on a published timetable that has itself been the subject of amendment proposals — check the current schedule against the primary text rather than against a summary, including ours.

NIST AI RMF

The US National Institute of Standards and Technology's AI Risk Management Framework is voluntary, and it is nonetheless the vocabulary a great many US enterprise and public-sector buyers now use in their questionnaires. It organises AI risk work into four functions — Govern, Map, Measure, Manage — and it is useful in an audit precisely because it is not a checklist: it asks whether you have a process, whether the process produces evidence, and whether anyone acts on what the evidence says. We walk through the four functions in plain English separately.

ISO/IEC 42001

ISO/IEC 42001:2023 is the management-system standard for AI — the AI analogue of ISO 27001 for information security. It is about the machinery around your AI rather than any individual model: who owns AI risk, how systems enter and leave the inventory, how decisions get documented, how the whole thing is reviewed. It is the standard most often asked for by counterparties who want to know that your AI governance will still exist next quarter.

Sector rules sit on top of these, never instead of them. A US health system deploying a triage assistant is answering to HIPAA and to its state's AI rules and to whichever of the above its buyers ask about. A lender is answering to fair-lending law and model-risk supervision. Those are the subjects of the domain guides linked at the foot of this piece.

What an AI audit is not

Three things, stated plainly, because each one gets sold as an audit somewhere.

It is not a certification. Certification is an accredited body attesting that a management system conforms to a standard, on a defined scope, for a defined period, under surveillance. An audit is a review and a report. Anyone offering to make you "EU AI Act certified" is describing something that does not exist in that form.

It is not a guarantee of compliance. An audit describes a system as it was, on the evidence available, at a point in time. Models drift, traffic changes, vendors ship, and rules move. A clean report is a strong record of diligence and a bad basis for complacency.

It is not a penetration test with a new label. Adversarial and security testing is one component — and a real one; we test with adversarial inputs rather than the vendor's demo set — but a security review that never asks who the system disadvantages, or whether anyone can explain a decision to the person it was made about, has not audited the AI. It has audited the software.

Independent review versus internal review

Internal review is necessary and it is not the same thing. The difference is not competence — internal teams usually know the system far better than any outsider will. The difference is the evidence chain and the incentive.

An internal review generally reports, directly or eventually, to the people who chose or built the system, and it tends to take that team's documentation as its evidence base. An independent review starts from the position that documentation is a claim. It is also the version a third party can rely on: a partner bank, an enterprise procurement team or a regulator asking "who checked this?" is asking specifically whether the checker had a stake in the answer. What "independent" actually means is worth its own read.

What you get at the end

An audit that produces a score and nothing else has told you very little. The deliverable that matters has four parts:

  • A scope statement — what was reviewed, what was not, and why. The exclusions are as important as the inclusions, and a report that has none is hiding something.
  • Findings, ranked. Not a flat list. A finding that creates real regulatory or legal exposure and a finding that is good practice deferred are different objects, and mixing them wastes the remediation budget on the wrong things. We describe how we grade a blocker against a recommendation separately.
  • Evidence. Each finding traceable to what produced it — a test output, a document, an interview, an absence. A finding you cannot trace is an assertion.
  • A fix path. Ordered by exposure, with an honest note on what is a configuration change and what is a rebuild.

And a named human auditor. We name ours, because a report nobody signs is a report nobody stands behind.

When an audit is worth doing

Honestly: not always, and not always yet. If your AI drafts internal text and touches no personal data and decides nothing about anyone, a full audit is probably not the best use of the budget this quarter — write down what it does, who owns it, and what would have to change for that answer to change.

It becomes worth doing when one of four things is true. Your AI influences a consequential decision about a person. A counterparty — a partner bank, an enterprise buyer, an insurer, a regulator — has started asking. You are relying on a vendor's compliance claims and have never tested one. Or you are about to scale something that currently works because a small team is watching it closely.

iDharma publishes three tiers, at fixed prices, with published turnarounds: a $5,000 Quick Scan at around a week, a $15,000 Compliance Audit at around two to three weeks, and a $25,000 Risk Audit at around four. What sits behind all three is the published verification methodology — the standard our own work is measured against, which you can read before you talk to us.

Frequently asked questions

What is an AI audit?
An AI audit is an independent, evidence-based review of an AI system measured against a named standard. It tests what the system does in operation rather than what its documentation claims, and produces a written record: the scope reviewed, the evidence examined, findings ranked by severity, and a path to fixing them.
Is an AI audit the same as a certification?
No. A certification is an accredited body attesting that a management system meets a standard, on a defined scope and for a defined period. An audit is a review and a report. iDharma does not issue certifications, and an audit report should never be read as a guarantee that a system is compliant.
How is an AI audit different from an internal model review?
The evidence chain and the incentive. An internal review is generally run by, or reports into, the team that built or bought the system, and tends to rely on that team's own documentation. An independent audit is run by a party with no stake in the result and treats vendor documentation as a claim to be tested, not as evidence.
Which standards is an AI system audited against?
It depends on where the system runs and what it decides. The three that apply most broadly are the EU AI Act, which is law for AI touching EU users; the NIST AI Risk Management Framework, which is voluntary but widely referenced; and ISO/IEC 42001, the management-system standard for AI. Sector rules — health, credit, insurance — sit on top of those, not instead of them.
How long does an AI audit take?
iDharma's published turnarounds run from about a week for a Quick Scan to about four weeks for a full Risk Audit, measured from approved scope to delivered report. The variable that moves it most is how quickly the organisation can produce evidence, which is why the document list is sent before the engagement starts.

Where your AI stands

Wondering where your AI stands?