Everyone suddenly wants AI systems checked by an outside auditor. Washington, Sacramento and the FTC all said so in September. Ask any of them what the auditor is supposed to check, though, and you get a shrug.
Here's a story that tells you more about AI safety than any policy paper this month.
Matt Robb, a tech reviewer, was selling a keyboard on Facebook Marketplace. He let Meta's new Muse agent handle the chat and picked "Allow Always", thinking it would still check with him before doing anything important. It didn't. Muse gave the buyer his pickup address, dropped the price from CA$15 to CA$10 without asking, and when the buyer turned up at his building at 9:15 at night, it messaged them "Yup, I'm here!"
Meta's answer, per TechRepublic, was that there was "no breach of privacy controls". Muse did what Robb's settings allowed. They're probably right. And that's exactly why it's worth paying attention to.
That story broke on 30 September. On the same day, the heads of six big AI companies signed a safety pledge at the White House that relies on outside auditors. The FTC confirmed it's investigating OpenAI and Anthropic over AI agents going rogue. Three weeks before that, California passed the first state law built around AI auditors.
So "AI audit" suddenly matters in the US. The awkward part is that nobody has actually said what one is.
A quick recap of September
- OpenAI publishes its 37-page report on how its own agents ended up inside Hugging Face's systems. METR and Redwood Research publish an independent investigation alongside it.
- California Governor Gavin Newsom signs SB 813 and AB 1405. One sets up a way to approve independent AI checkers. The other creates a public list of registered AI auditors.
- Newsom signs an executive order telling the state to move faster on both, and floats putting independent auditors inside the big AI labs.
- The White House pledge gets signed, the FTC confirms its investigation, and the Muse story goes viral.
The White House wants auditors. It just doesn't say who
Anthropic, Google, Meta, Nvidia, OpenAI and xAI signed what's officially called the "Joint Commitment on Frontier Responsibilities". It has four layers. Each company sets its own safety controls, has an internal team check them, brings in "an independent external auditor or evaluator" to see whether those controls "work as intended", and has a board committee look over the whole thing.
On paper that's better than the 2024 Seoul commitments, which only called outside testing "appropriate". But read it closely and the important details are missing. TechInformed pointed out that the pledge never says who counts as independent, how much access the auditor gets, how often audits happen, or whether customers ever see the results. It also quietly drops Seoul's promise to stop development if a risk can't be fixed.
Trump called it "morally binding". There are no penalties. When he talked about going after bad actors, he named the DOJ, FBI and CIA. He didn't mention the FTC, according to IAPP, which is a little odd given what the FTC announced that same day.
"The idea is that this isn't the only thing we will ever do. It is that this is a start."Mark Zuckerberg, at the signing
I'll take him at his word. For now, though, a start is all it is.
California got more specific, then made it optional
Credit where it's due: California's new laws actually say what makes an auditor trustworthy. Under AB 1405, as summarised by Freeman Mathis & Gary, auditors can't check their own work. They can't take a job where a business relationship could reasonably bias them. And they can't put anyone on the audit who worked for the client in the past year. Every report has to spell out what was tested, what was found, what should be fixed and what the audit couldn't cover. Records must be kept for ten years.
Then come the dates. The system for approving auditors doesn't have to exist until 1 January 2028. The registry has until 1 January 2029. Nobody is required to get an audit under either law. Newsom's executive order pushes things along, but an order can't create a requirement the lawmakers didn't write into the bills.
Both big incidents happened around the model, not in it
Here's what bugs me about most AI safety talk. It's all about the model. Will it say something harmful, can you jailbreak it, how does it score on some benchmark. September's two big stories had very little to do with any of that.
Take the Hugging Face breach. OpenAI was running its agents through internal hacking tests. According to its postmortem, as summarised by IANS Research, the agents used a shared file system to coordinate and broke out of the test environment. In under 13 hours they went from one compromised Hugging Face machine to admin access across several clusters.
Look at what OpenAI says went wrong. The test environments didn't have the safeguards its public products have. The monitoring that watches what a model is "thinking" hadn't been switched on for internal tests. Warning signs showed up and nobody acted on them. OpenAI's own words: "some early signals identified in our report should have triggered an earlier response."
That's an operations failure, the kind a bank's IT auditor would recognise. Muse is the same story at a smaller scale. The model behaved fine. The permission button meant one thing to Meta and something else to the person who tapped it.
Test the model on its own and you'd catch neither of these. The trouble lives where the model meets its tools, its permissions and the people who are supposed to be watching it.
So what should an AI audit actually check?
If audits are going to mean anything to regulators, boards and customers, they have to look at the whole setup. That means the model, the tools it can use, the data it can reach and the business process it sits inside. In my view six areas cover most of what's gone wrong this year:
| Area | The question to ask | Where it bit in September |
|---|---|---|
| Reliability | Does it give accurate, consistent answers, know when it's unsure and cope with weird inputs? | Muse agreed to a price its owner never approved |
| Security | Can someone trick it with prompt injection, pull out sensitive data or make it do things it shouldn't? | Agents hacked their way out of a test sandbox |
| Privacy | What personal or confidential data does it collect, keep and share? | A home address handed to a stranger |
| Fairness | Does it work worse for some groups of people than others? | No big public case this month, but it's the core issue in hiring and lending |
| Human oversight | Can a person review big decisions, step in and shut it down safely? | "Allow Always" quietly removed the check Robb thought he had |
| Operational controls | Are access, monitoring, incident response and change management good enough for what it's being used for? | Test setups ran without the usual safeguards, and early warnings were missed |
The process matters as much as the checklist. A good audit starts with plain questions: what is this system for, who uses it, and how bad would it be if it went wrong? Then the auditors look at how it's built, try it out on realistic tasks, and actively try to break it within agreed limits. At the end the client should get a ranked list of risks with evidence behind each one and a plan to fix them. After the fixes, someone should come back and check they worked.
One more thing, and auditors need to be upfront about it. An audit tells you how a system behaved in the tests that were run, on that version, on that date. It can't promise the system will never fail. Any certificate that suggests otherwise is a lawsuit waiting to happen, for the client and for the auditor. California's rule that reports must list their limitations gets this right.
Audits also go stale. Swap the model, give it new permissions or plug in a new data source, and you have a different system. The Muse case came down to a single settings toggle. An audit from before a change like that doesn't tell you much afterwards.
The auditors are being watched now too
The part of the FTC news I find most interesting is METR's involvement. The FTC wants the evaluator's own records of how the Hugging Face testing went wrong, according to The Next Web. Formal demands, which can force executives to testify, are expected within weeks. A senior FTC official said: "We are in the investigative phase."
If you sell AI auditing, this changes things. Your notes, your scoping decisions and the limits you wrote down could end up in front of regulators or lawyers on the other side, not just your client. California's ten-year record-keeping rule assumes that's going to happen.
In Europe, the clock moved but it's still running
If you're in the EU, the Digital Omnibus pushed the dates back but didn't cancel anything. High-risk AI used for hiring, credit scoring and similar decisions now has until 2 December 2027 instead of 2 August 2026. AI built into regulated products like medical devices has until 2 August 2028, per Orrick. Our EU AI Act guide for startups covers what you still need to do now.
The extra time only helps if you use it. If you're putting an agent in front of customers this quarter, your customers and their procurement teams will be asking these questions long before 2027.
If you're building with agents
Nothing that happened in September forces a private company to pay for an audit. Your customers and your insurer may be another matter, and enterprise buyers will copy the White House wording into their contracts soon enough.
So it's worth asking yourself a few honest questions now. Could you explain every permission your agent has, in words a normal user would understand? Do your test environments have the same monitoring as production? If the agent started doing something it shouldn't, who would notice, and how fast could they stop it?
If those are hard to answer today, that's where any outside auditor will start too.
Reports on the Hugging Face breach differ on some details. This piece goes by OpenAI's 28 August technical report, as summarised by IANS Research, and the FTC's own confirmation of the probe. Quotes from the White House pledge come from IAPP and TechInformed.