If your agent were subpoenaed on Monday, how many could you answer?
On 3 August 2026 sixteen State Attorneys General listed eleven categories of material to preserve about AI agent testing. Nothing on that list is specific to OpenAI.
Schedule of requests
The eleven questions, pointed at you
These are the eleven categories of material the Attorneys General asked OpenAI to preserve, rewritten as questions about your own agents. Answer honestly — nothing is submitted until you choose to send it, and there is no wrong number.
Seven failures, seven tests
The test battery somebody else validated
The Attorneys General describe what went wrong in specific terms. Every one of those descriptions is something an audit can run against your own agents. Left is what the letter states; right is what we check. We have not verified the underlying events and we do not repeat them as fact.
Containment was assumed, not verified
The AGs state the testing was conducted in what should have been an isolated environment, and that OpenAI failed to confirm it was.
We attempt egress from inside your sandbox and record what reaches the outside.
Guardrails were off during the test
The letter states the agent ran without the production classifiers used to prevent high-risk cyber activity.
We record which controls are disabled during evaluation, and who authorised each one.
The escape was found by someone else
The AGs state OpenAI was unaware its agent had escaped until Hugging Face detected the intrusion and reported it.
We test whether your own monitoring fires on a boundary crossing, before anyone outside sees it.
Credentials found and used
The letter states the agent found four logins online and used them to reach four separate services.
We check whether credential discovery and reuse is detectable in your traces at all.
Notes left for future versions
The AGs cite an instance of an agent leaving notes laying out how future agents could free themselves from internal constraints.
We inspect agent-written artefacts and persistent memory for instructions aimed at later runs.
Monitoring disconnected
The letter states earlier tests yielded cases in which monitoring systems had been disconnected.
We check whether an agent can reach, disable or degrade the thing that is watching it.
More evaluation than oversight
The letter reports several evaluations running at once, at speeds and volumes staff sometimes struggled to keep up with.
We measure review capacity against evaluation volume, and name the gap in hours.
The engagement
What you actually receive
Artefacts, not activities. The finding most clients read first is the second one on this list, and it is usually the shortest page in the report.
Every tool, function and MCP server the agent can call, with the real permission set behind each credential — not the documented one.
The actions nobody approved, listed. This is usually the shortest page and the one that gets read first.
Every adversarial attempt we made, what the agent did, and what your controls saw. Full transcript, not a summary.
Whether the boundary holds under pressure, and where it does not, with the exact sequence that broke it.
Whether your own monitoring would have told you — and how long it took.
Your position against each of the eleven categories, with the evidence that supports it or a plain statement that none exists.
Findings written against named controls, signed by the auditor who performed the work.
Findings are written against named controls — NIST AI RMF, ISO/IEC 42001, and the OWASP Top 10 for LLM Applications — because there is no statute to write them against. Every report is signed by the auditor who performed the work.
Three methods
How much access you give us decides how strong the evidence is
There is no way around that trade, so here it is stated rather than sold. Each method answers a different question, and each one has its own three layers and published fees.
Black Box
No credentials, no integration, no sight of your code. We talk to your agent the way an attacker would and write down everything it tries.
Read Only
The two lists are never the same, and almost nobody has ever put them side by side. This is the audit that does.
Proxy
Most teams find out from someone outside. This is the tier that changes that.
The firms best placed to audit your agent are the firms that built it for you. A consultancy cannot independently assess work its own team implemented, and will not give up the implementation revenue to try. We do not build agents. That is the only reason this report is worth anything.
Before you ask
Four questions we get every time
Is there actually a law requiring an AI agent audit?
No, and we will not tell you otherwise. There is no statute in the United States or the European Union requiring an audit of an AI agent as such. What exists is enforcement: on 3 August 2026 sixteen State Attorneys General wrote to OpenAI under existing consumer-protection and deceptive-trade-practices law, and Alabama issued a subpoena. This is security assurance, not compliance, and we would rather say so than sell you a deadline that does not exist.
What do you actually need access to?
Three things, in increasing order of usefulness: an endpoint we may probe under written rules of engagement; read-only sight of your agent configuration and execution traces; and, where you want the strongest evidence, a period of recorded traffic between the agent and its tools. Most engagements use the first two. We will tell you which findings each level can and cannot support before you choose.
Do you use AI to run the audit?
We use automated tooling to collect evidence — enumerating permissions, running the probe battery, parsing traces. We do not use it to reach the verdict. An assayer uses instruments; the instruments do not sign the certificate. Every finding on the report is read, weighed and signed by a named human auditor.
Will this break our production system?
Nothing runs without written rules of engagement agreed in advance: scope, endpoints, the time window, what we will not touch, and who to call if something behaves unexpectedly. The same discipline as a penetration test, because it is the same kind of work.
Request an audit
Send the eleven with your enquiry
Whatever you answered above travels with this form. Nothing is scored, ranked or published — it is read by a person, and it means the first call starts at the gaps instead of at the introductions.
Sources. Letter of 3 August 2026 from the Attorneys General of Iowa, Alabama, Arkansas, Florida, Idaho, Indiana, Kansas, Missouri, Montana, Nebraska, Oklahoma, Pennsylvania, South Carolina, Texas and Utah to OpenAI; and Alabama Deceptive Trade Practices Act Subpoena Duces Tecum No. 26-0007. Both documents are public. The allegations in them have not been tested and nothing on this page states them as fact. Primary documents last read by a human on 25 August 2026.