AI AGENT ASSURANCE · INDEPENDENT · NO STATUTE REQUIRED

If your agent were subpoenaed, how many could you answer?

Sixteen Attorneys General listed eleven categories of material to preserve about agent testing - and nothing on it is specific to OpenAI.


A reviewer with silver hair, in a navy blazer, a closed notebook at the edge of the desk, seated at a dark stone desk by a window in a warm, low-lit office, signing a printed page with further papers and a stoneware cup beside them.
Independent means no stake in the answer
Capability boundary Containment Detection Credential use Evidence record

Our promise

“A system card is a claim. The trace is evidence.”

Every finding is written against a named control — defensible line by line, to anyone who asks. The fee starts at $8,500, and nothing is charged until you approve it.

Request this agent audit
The case file

Agent assurance, in three chapters

The Letter

On 3 August 2026 the Attorneys General of sixteen states wrote jointly to OpenAI, stating it may have violated consumer-protection law in its testing of autonomous agents, and listing eleven categories of material to preserve. Alabama issued a subpoena the same month.

The Gap

Nothing on that list is specific to OpenAI. Any organisation running agents either holds a per-action record, evidence that it detected the breach itself, and a dated safety policy — or it does not. Most hold logs. Almost nobody has assembled them into such a record.

The Office

We are the independent auditor no law has yet required. iDharma probes the agent itself, reconciles what it may reach against what it was granted, tests whether your own monitoring notices at all — and writes your position against all eleven categories, signed by the auditor who ran it.

What is an agent?

Software that acts before you do.

A system that plans its own steps, calls tools and acts without a person between one instruction and the next.

Who this is for № 01
  • Teams running agents against production systems
  • Anyone whose agent holds credentials of its own
  • Platforms exposing tools or MCP servers to agents
  • Risk owners asked to sign off on an agent’s release
Agent audit · iDharma · Presented for assay
When to run one № 02
  • Before an agent is issued a credential of its own
  • Before first use against anything customer-facing
  • Whenever a tool, a scope or the model changes under it
  • After any incident somebody outside reported to you
Agent audit · iDharma · Presented for assay
Who has to answer

The agent is yours. The model is theirs.

You

Whoever deployed the agent

The eleven categories are about a deployment, not a model. What it was allowed to reach, what it did reach, who authorised the test and who knew afterwards are all facts about your environment. Nobody else holds them, and nobody else can produce any of them on your behalf.

Your provider

The people who built the model

Publishes system cards and safety evaluations, and those are often genuinely good work done by serious people. But they describe a model under their own conditions, not your agent under yours. The difficulty is not their quality - it is that the thing they tested is not the thing you shipped.

The catch

What their evaluation covers

Only for the model itself, and only as it was configured on their side. The moment you attach a tool, issue a credential or widen a scope, you have built something that nobody anywhere has ever evaluated. If your agent can reach a system their evaluation never saw, it does not cover you and it never did.

What most teams assume

“Our provider tested the model, so we’re covered.”

What the letter asks

Every question is about your deployment, not theirs.

It is the first gap we find on most engagements.

  • Who it is for
  • AI product teams
  • Platform & infrastructure
  • Security & risk
  • Legal & compliance
  • Agent and MCP tool vendors
Why this matters now
Hands at a dark desk under a low lamp, one holding a fountain pen against a printed report of tabulated figures and bar charts, the other resting flat on the page: the moment an account of what happened is written down rather than remembered.
01 Enforcement arrived before legislation. There is no statute to comply with and a subpoena already exists. That is an unusual order of events, and it is the one we are in.
02

The eleven ask for records, not assurances. A summary of an incident is not a record of one, and a summary is what most teams hold.

03

Discovery by somebody outside is itself a finding rather than a footnote to one. In the case the letter describes, that is exactly how it happened.

04

Retention is the quiet one. Traces roll off on a schedule set long before anybody thought to ask for them, and that schedule was never a decision.

The 60-second check

Three questions. Then you’ll know.

No email, no signup. A starting point, not a determination.

0 of 3

What it does -

A chatbot answers. An agent acts. The line is whether anything happens between one instruction and the next without a person in it.

What it reaches -

The blast radius is the permission scope, not the prompt. What it can reach is the question. What it usually does is not.

The record -

A summary is not a record. Per action: what it did, what it touched, and which model version was running when it did.

Going further: the eleven a preservation letter actually asked for.

Schedule of requests

Eleven questions. Pointed at you.

No email. No signup. Nothing is sent until you choose to send it.

Where the eleven come from

Not our list. On 3 August 2026 the Attorneys General of sixteen states wrote to OpenAI and named eleven categories of material to preserve. Each one is turned round here — from preserve this into could you produce this — and nothing else about it is changed. Nothing on the list is specific to OpenAI: an organisation running agents either holds these records or it does not.

The allegations behind the letter have not been tested, and nothing here states them as fact. It is used as what it is — a published account of what somebody with subpoena power thought worth asking for.

0 of 11

The incident record -

What a no means. No per-action trace. A summary is not a record.

How you found out -

What a no means. No independent detection. Discovery depends on a third party.

Which model ran -

What a no means. Model provenance not recorded per run.

Your own review -

What a no means. No written review, or a review that diverges from public statements.

Credential use -

What a no means. Credential discovery and reuse is untracked.

Offensive evaluations -

What a no means. Offensive testing runs without a named authoriser.

Prior incidents -

What a no means. No incident history. "None that we know of" is not an answer.

Self-persistence -

What a no means. Agent-written artefacts are not inspected.

Evaluation safety policy -

What a no means. No dated policy. A policy written after the fact proves nothing.

Concerns raised -

What a no means. Internal concerns are discoverable and unmanaged. Usually the sharpest exposure.

Who knew -

What a no means. No named custodians.

The four moments

Four moments, and two happen before you ask.

Each of these is measured against the moment an agent crosses a line - so they cannot be added up, and the first two are usually decided long in advance.

  1. Detect

    Hour one

    Your own monitoring, not word from outside. The gap between crossing a line and noticing is the measure.

  2. Preserve

    Before rotation

    Traces, prompts and tool calls held before retention rolls them off. Overwritten evidence is not recoverable.

  3. Reconstruct

    Days, not weeks

    An account of what ran, under whose authority and what it reached. Assembled once, it answers again.

  4. Re-test

    Every release

    A new tool, a new scope or a new model is a new agent. The last audit describes something you no longer run.

The trap

Preserve sits in front of reconstruct, and nobody schedules it. Your retention policy was set months before any of this - so by the time a question arrives, the evidence that would answer it has either been kept or has already rolled off.

The ask & the artefact

What the letter asks, what we ship

12 things to be answerable for, and the artefact that answers each. Paired, so every claim on this page can be checked against the ask beside it.

The complete incident record Every action, every service touched
A per-action trace rebuilt from your own logs, or a written finding that it cannot be.
How you found out You, or somebody outside
A detection test: we cross a boundary and record whether your monitoring said so first.
Which model was running Version and build, per run
Model and version bound to each execution, so any trace names what actually produced it.
Containment Verified, not assumed
Egress attempted from inside your sandbox, with everything that reached the outside logged.
Guardrails during testing What was off, and who allowed it
The controls disabled for each evaluation, named, with the person who authorised each one.
Credential discovery and reuse Account-level access never granted
Whether an agent finding and using a credential is visible in your traces at all.
Agent-written artefacts Notes left for later runs
Persistent memory and written output inspected for instructions aimed at future versions.
Monitoring integrity Can it reach what watches it
A test of whether the agent can disable, degrade or blind the thing that is monitoring it.
Offensive evaluations A named authoriser for each
The register of evaluations that prompt toward exploitation, and who signed off on each.
Evaluation safety policy Dated before the test, not after
Your policy read against your test dates, so one written afterwards is visible as one.
Review capacity Oversight against test volume
Human review measured against evaluation throughput, with the shortfall named in hours.
Named custodians Who knew, and who holds it
The people who can answer for each artefact, named, so a request has somewhere to land.
The engagement

Autonomous agents, independently audited

From a single tool-using assistant to a fleet in production.

  1. Intake

    Which agents are in scope, and what each one can actually reach.

  2. Test

    Probes, permissions reconciled, containment and detection tested.

  3. Sign off and hand over

    You see the draft first. Then the report, transcript and answers.

Request your bias audit
An auditor in a charcoal suit and open-collared white shirt, standing against a warm pale wall and pointing into the open space alongside.
The numbers are not negotiable - that is what you are buying.
Struck in your favour

Why teams choose iDharma to audit their agents

Genuinely independent

We build, sell and operate no AI agents, and we take no fee that depends on what we find.

Written to named controls

Every finding maps to NIST AI RMF, ISO/IEC 42001 or the OWASP LLM Top 10 - never to a statute.

The transcript, not a summary

You get every probe we ran and everything the agent attempted, in full, rather than a digest.

Signed by a named person

One auditor reads the evidence, reaches the finding, and can be asked about it afterwards.

Four marks, struck on every report.

Deliverables

What you get

Concrete artefacts, each with a name and a format - you know what lands before you buy.

Signed audit report

The full engagement: method, scope, every probe and its result, and the findings in plain language - each written against a named control from NIST AI RMF, ISO/IEC 42001 or the OWASP Top 10 for LLM Applications rather than against a statute, ranked by consequence, and signed by the auditor who performed the work.

Register

Capability register

Every tool, function and MCP server the agent can call, with the real permission set behind each credential rather than the documented one.

Shortlist

Declared-versus-permitted gap

The actions nobody approved but the credentials allow, listed. Usually the shortest page in the report and reliably the first one read.

Full log

Probe transcript

Every adversarial attempt we made, what the agent did in response, and what your own controls saw. The whole transcript, not a summary of it.

Finding

Containment finding

Whether the boundary holds under pressure and where it does not, with the exact sequence that broke it written out so you can run it again.

Finding

Detection finding

Whether your own monitoring would have told you, what it recorded when the boundary moved, and how long it took to say anything at all.

Position

The eleven answers

Your position against each of the eleven categories, with the evidence behind it - or a plain statement, in writing, that no evidence exists.

Format & fee

Real numbers, upfront.

Scope
Set by the method you choose
Access
Endpoint, configuration or proxy
Re-assay
Every twelve months - $7,200 against your known baseline

No law fixes the scope, so the fee follows the method - published per tier, nothing charged until you approve.

Request an agent audit
Black Box · One agent $8,500 one-off
  • Adversarial probe battery, run to plan
  • Full transcript, not a summary of it
  • Rules of engagement agreed first
  • Findings signed by a named auditor
Show your hand

Four things you have to be able to produce

Nobody grades your intentions. Each of these is either in your hand on the day somebody asks, or it is not.

The record,
per action

A trace of what the agent actually did, action by action, including every other account and service it touched. A summary of an incident is not a record of one at all.

The detection,
your own

Evidence that you found it rather than somebody outside. Discovery by a third party is the finding itself, not a footnote to it - and it is the ordinary case, not the rare one.

The authority,
dated

A policy governing evaluation safety, dated before the test it governs, and a named person who authorised each run against it. One written afterwards proves nothing.

The people,
named

The people involved in what happened or aware of it, named. Ten of the eleven categories ask for documents; this one asks for people and nobody has it ready.

Four cards, and the date on each one is part of the card.

FAQ

Plain answers

Whether there is a law, what we need, and what we will not do.

Request an agent audit
Is there actually a law requiring an AI agent audit?

No, and we will not tell you otherwise. No statute in the US or the EU requires an audit of an AI agent as such. What exists is enforcement: sixteen State Attorneys General wrote to OpenAI on 3 August 2026, and Alabama issued a subpoena.

Our model provider publishes safety evaluations. Are we covered?

For the model, and only as configured on their side. The moment you attach a tool, issue a credential or widen a scope you have built something nobody has evaluated - and that is what the eleven categories ask about.

What do you actually need access to?

Three things, in increasing order of usefulness: an endpoint we may probe under written rules of engagement; read-only sight of configuration and traces; and, for the strongest evidence, a period of recorded traffic between the agent and its tools.

Do you use AI to run the audit?

To collect evidence, yes - enumerating permissions, running the probe battery, parsing traces. Not to reach the verdict. An assayer uses instruments; the instruments do not sign the certificate.

Will this break our production system?

Nothing runs without written rules of engagement agreed first: scope, endpoints, the time window, what we will not touch, and who to call if something behaves unexpectedly. The same discipline as a penetration test.

Request an audit

Send the eleven with your enquiry

Whatever you answered above travels with this form. Nothing is scored, ranked or published.

What we need from you

Nothing you do not already have. Most of this is what your own team knows about the agents you run, and we name what we need in writing before you commit to anything.

  1. Which agents you run, and what each one decides
  2. How many there are, and whether any are in production
  3. What each can reach - tools, APIs, the data behind
  4. Whether it acts under its own identity or a person's
  5. Your target date for a read, if you have one

What happens next

  1. You send this, with whatever you answered above.
  2. We read it and reply within two working days.
  3. Nothing is charged until you approve the scope.

Your answers to the eleven will be attached.

Request an agent audit

Send your enquiry

Six fields and two boxes, only three of them required, and nothing to attach.

Your answers to the eleven will be attached.

Sources & standing

Where this page gets its facts

Where the claims on this page come from, and what they are worth - stated, not assumed.

What it is drawn from

  • Letter of sixteen State Attorneys General to OpenAI
  • Alabama subpoena duces tecum No. 26-0007
Letter dated
3 August 2026
Last read
25 August 2026

What it means

  • Both documents are public. The allegations in them have not been tested, and nothing here states one as fact.
  • No statute requires an AI agent audit. This is security assurance, and the page says so wherever it could be misread.

Scope & limitation

  • Findings are written against NIST AI RMF, ISO/IEC 42001 and the OWASP Top 10 for LLM Applications.
  • General information, not legal advice. Engage counsel before resting a binding decision on any of it.

Something on this page out of date?

Tell us