On 3 August 2026 the Attorneys General of sixteen states wrote jointly to OpenAI, stating it may have violated consumer-protection law in its testing of autonomous agents, and listing eleven categories of material to preserve. Alabama issued a subpoena the same month.
If your agent were subpoenaed, how many could you answer?
Sixteen Attorneys General listed eleven categories of material to preserve about agent testing - and nothing on it is specific to OpenAI.
Our promise
“A system card is a claim. The trace is evidence.”
Every finding is written against a named control — defensible line by line, to anyone who asks. The fee starts at $8,500, and nothing is charged until you approve it.
Request this agent auditAgent assurance, in three chapters
Nothing on that list is specific to OpenAI. Any organisation running agents either holds a per-action record, evidence that it detected the breach itself, and a dated safety policy — or it does not. Most hold logs. Almost nobody has assembled them into such a record.
We are the independent auditor no law has yet required. iDharma probes the agent itself, reconciles what it may reach against what it was granted, tests whether your own monitoring notices at all — and writes your position against all eleven categories, signed by the auditor who ran it.
Software that acts before you do.
A system that plans its own steps, calls tools and acts without a person between one instruction and the next.
- Teams running agents against production systems
- Anyone whose agent holds credentials of its own
- Platforms exposing tools or MCP servers to agents
- Risk owners asked to sign off on an agent’s release
- Before an agent is issued a credential of its own
- Before first use against anything customer-facing
- Whenever a tool, a scope or the model changes under it
- After any incident somebody outside reported to you
The agent is yours. The model is theirs.
Whoever deployed the agent
The eleven categories are about a deployment, not a model. What it was allowed to reach, what it did reach, who authorised the test and who knew afterwards are all facts about your environment. Nobody else holds them, and nobody else can produce any of them on your behalf.
The people who built the model
Publishes system cards and safety evaluations, and those are often genuinely good work done by serious people. But they describe a model under their own conditions, not your agent under yours. The difficulty is not their quality - it is that the thing they tested is not the thing you shipped.
What their evaluation covers
Only for the model itself, and only as it was configured on their side. The moment you attach a tool, issue a credential or widen a scope, you have built something that nobody anywhere has ever evaluated. If your agent can reach a system their evaluation never saw, it does not cover you and it never did.
“Our provider tested the model, so we’re covered.”
Every question is about your deployment, not theirs.
It is the first gap we find on most engagements.
- Who it is for
- AI product teams
- Platform & infrastructure
- Security & risk
- Legal & compliance
- Agent and MCP tool vendors
The eleven ask for records, not assurances. A summary of an incident is not a record of one, and a summary is what most teams hold.
Discovery by somebody outside is itself a finding rather than a footnote to one. In the case the letter describes, that is exactly how it happened.
Retention is the quiet one. Traces roll off on a schedule set long before anybody thought to ask for them, and that schedule was never a decision.
Three questions. Then you’ll know.
No email, no signup. A starting point, not a determination.
Going further: the eleven a preservation letter actually asked for.
Your 60-second check
Eleven questions. Pointed at you.
No email. No signup. Nothing is sent until you choose to send it.
Not our list. On 3 August 2026 the Attorneys General of sixteen states wrote to OpenAI and named eleven categories of material to preserve. Each one is turned round here — from preserve this into could you produce this — and nothing else about it is changed. Nothing on the list is specific to OpenAI: an organisation running agents either holds these records or it does not.
The allegations behind the letter have not been tested, and nothing here states them as fact. It is used as what it is — a published account of what somebody with subpoena power thought worth asking for.
0 of 11
Your evidence check
Eleven questions
Nothing is saved and nothing is sent. No account, no login — answer as many as you like and close it.
Request an agent audit
Any of the eleven you answered travel with this. Nothing is sent until you press send.
Four moments, and two happen before you ask.
Each of these is measured against the moment an agent crosses a line - so they cannot be added up, and the first two are usually decided long in advance.
-
Detect
Hour oneYour own monitoring, not word from outside. The gap between crossing a line and noticing is the measure.
-
Preserve
Before rotationTraces, prompts and tool calls held before retention rolls them off. Overwritten evidence is not recoverable.
-
Reconstruct
Days, not weeksAn account of what ran, under whose authority and what it reached. Assembled once, it answers again.
-
Re-test
Every releaseA new tool, a new scope or a new model is a new agent. The last audit describes something you no longer run.
Preserve sits in front of reconstruct, and nobody schedules it. Your retention policy was set months before any of this - so by the time a question arrives, the evidence that would answer it has either been kept or has already rolled off.
What the letter asks, what we ship
12 things to be answerable for, and the artefact that answers each. Paired, so every claim on this page can be checked against the ask beside it.
- The complete incident record Every action, every service touched
- A per-action trace rebuilt from your own logs, or a written finding that it cannot be.
- How you found out You, or somebody outside
- A detection test: we cross a boundary and record whether your monitoring said so first.
- Which model was running Version and build, per run
- Model and version bound to each execution, so any trace names what actually produced it.
- Containment Verified, not assumed
- Egress attempted from inside your sandbox, with everything that reached the outside logged.
- Guardrails during testing What was off, and who allowed it
- The controls disabled for each evaluation, named, with the person who authorised each one.
- Credential discovery and reuse Account-level access never granted
- Whether an agent finding and using a credential is visible in your traces at all.
- Agent-written artefacts Notes left for later runs
- Persistent memory and written output inspected for instructions aimed at future versions.
- Monitoring integrity Can it reach what watches it
- A test of whether the agent can disable, degrade or blind the thing that is monitoring it.
- Offensive evaluations A named authoriser for each
- The register of evaluations that prompt toward exploitation, and who signed off on each.
- Evaluation safety policy Dated before the test, not after
- Your policy read against your test dates, so one written afterwards is visible as one.
- Review capacity Oversight against test volume
- Human review measured against evaluation throughput, with the shortfall named in hours.
- Named custodians Who knew, and who holds it
- The people who can answer for each artefact, named, so a request has somewhere to land.
Autonomous agents, independently audited
From a single tool-using assistant to a fleet in production.
-
Intake
Which agents are in scope, and what each one can actually reach.
-
Test
Probes, permissions reconciled, containment and detection tested.
-
Sign off and hand over
You see the draft first. Then the report, transcript and answers.
Why teams choose iDharma to audit their agents
Genuinely independent
We build, sell and operate no AI agents, and we take no fee that depends on what we find.
Written to named controls
Every finding maps to NIST AI RMF, ISO/IEC 42001 or the OWASP LLM Top 10 - never to a statute.
The transcript, not a summary
You get every probe we ran and everything the agent attempted, in full, rather than a digest.
Signed by a named person
One auditor reads the evidence, reaches the finding, and can be asked about it afterwards.
Four marks, struck on every report.
What you get
Concrete artefacts, each with a name and a format - you know what lands before you buy.
Signed audit report
The full engagement: method, scope, every probe and its result, and the findings in plain language - each written against a named control from NIST AI RMF, ISO/IEC 42001 or the OWASP Top 10 for LLM Applications rather than against a statute, ranked by consequence, and signed by the auditor who performed the work.
Capability register
Every tool, function and MCP server the agent can call, with the real permission set behind each credential rather than the documented one.
Declared-versus-permitted gap
The actions nobody approved but the credentials allow, listed. Usually the shortest page in the report and reliably the first one read.
Probe transcript
Every adversarial attempt we made, what the agent did in response, and what your own controls saw. The whole transcript, not a summary of it.
Containment finding
Whether the boundary holds under pressure and where it does not, with the exact sequence that broke it written out so you can run it again.
Detection finding
Whether your own monitoring would have told you, what it recorded when the boundary moved, and how long it took to say anything at all.
The eleven answers
Your position against each of the eleven categories, with the evidence behind it - or a plain statement, in writing, that no evidence exists.
Real numbers, upfront.
- Scope
- Set by the method you choose
- Access
- Endpoint, configuration or proxy
- Re-assay
- Every twelve months - $7,200 against your known baseline
No law fixes the scope, so the fee follows the method - published per tier, nothing charged until you approve.
Request an agent audit- Adversarial probe battery, run to plan
- Full transcript, not a summary of it
- Rules of engagement agreed first
- Findings signed by a named auditor
Four things you have to be able to produce
Nobody grades your intentions. Each of these is either in your hand on the day somebody asks, or it is not.
The record,
per action
A trace of what the agent actually did, action by action, including every other account and service it touched. A summary of an incident is not a record of one at all.
The detection,
your own
Evidence that you found it rather than somebody outside. Discovery by a third party is the finding itself, not a footnote to it - and it is the ordinary case, not the rare one.
The authority,
dated
A policy governing evaluation safety, dated before the test it governs, and a named person who authorised each run against it. One written afterwards proves nothing.
The people,
named
The people involved in what happened or aware of it, named. Ten of the eleven categories ask for documents; this one asks for people and nobody has it ready.
Four cards, and the date on each one is part of the card.
Plain answers
Whether there is a law, what we need, and what we will not do.
Request an agent auditIs there actually a law requiring an AI agent audit?
No, and we will not tell you otherwise. No statute in the US or the EU requires an audit of an AI agent as such. What exists is enforcement: sixteen State Attorneys General wrote to OpenAI on 3 August 2026, and Alabama issued a subpoena.
Our model provider publishes safety evaluations. Are we covered?
For the model, and only as configured on their side. The moment you attach a tool, issue a credential or widen a scope you have built something nobody has evaluated - and that is what the eleven categories ask about.
What do you actually need access to?
Three things, in increasing order of usefulness: an endpoint we may probe under written rules of engagement; read-only sight of configuration and traces; and, for the strongest evidence, a period of recorded traffic between the agent and its tools.
Do you use AI to run the audit?
To collect evidence, yes - enumerating permissions, running the probe battery, parsing traces. Not to reach the verdict. An assayer uses instruments; the instruments do not sign the certificate.
Will this break our production system?
Nothing runs without written rules of engagement agreed first: scope, endpoints, the time window, what we will not touch, and who to call if something behaves unexpectedly. The same discipline as a penetration test.
Send the eleven with your enquiry
Whatever you answered above travels with this form. Nothing is scored, ranked or published.
Where this page gets its facts
Where the claims on this page come from, and what they are worth - stated, not assumed.
What it is drawn from
- Letter of sixteen State Attorneys General to OpenAI
- Alabama subpoena duces tecum No. 26-0007
- Letter dated
- 3 August 2026
- Last read
- 25 August 2026
What it means
- Both documents are public. The allegations in them have not been tested, and nothing here states one as fact.
- No statute requires an AI agent audit. This is security assurance, and the page says so wherever it could be misread.
Scope & limitation
- Findings are written against NIST AI RMF, ISO/IEC 42001 and the OWASP Top 10 for LLM Applications.
- General information, not legal advice. Engage counsel before resting a binding decision on any of it.
Something on this page out of date?
Tell usFrom Insights
Before you commission one
How to Prepare for an AI Audit: The Readiness Checklist
Six things to have ready before the engagement starts. Assembling them takes a fortnight off the clock — and tends to find the first two findings before an auditor does.
Startups, Meet Your AI Stack: Budget‑Friendly Tools That Scale
For early-stage founders, building an AI-powered toolkit doesn’t have to break the bank. From ideation to growth mode, here’s how startups can tap into affordable, effective AI tools to autom
The Missing Link in Corporate AI Training
In today’s fast-paced AI era, companies are investing more in training—but why is most of it still missing the mark?
Colorado SB 26-189: Four Duties, and the One Nobody Budgets For
Notice before, disclosure after, human review on request, records for three years. Three are policy changes. The second is an engineering project, and it arrives late.