
The FRONTIER Act: What Independent AI Auditors Must Test
Yoni Fraimorice
OpenAI now supports a federal requirement for outside organizations to assess safety inside leading AI labs. That is a useful change. But what would those auditors actually test?
POLITICO's report, carried by Yahoo, says OpenAI supports the FRONTIER Act's provision for independent verification organizations, or IVOs. This is support for a specific mandate, not necessarily every part of the bill.
The hard question is whether an auditor can find dangerous behavior before it becomes another incident.
What the proposal actually requires
The FRONTIER Act is proposed legislation, not an audit system already in force. This article uses the bill text published by its sponsor. Its terms could change.
Section 5 would require "very large frontier developers" to retain a licensed IVO for ongoing assessment. The draft defines this tier using both revenue above $5 billion and AI development spending of at least $10 billion over the previous 36 months, including affiliates.
This is different from the annual audit duty for the separate "large" developer tier.
The IVO would assess the safety framework, governance, risk monitoring, and fixes. Crucially, this covers internal model use as well as public releases. Reports would go to the developer and the regulator at least every six months, with additional reporting when important findings change.
"Embedded" should therefore mean ongoing access, not just an office inside the lab. The draft requires timely access to necessary unredacted records, people, and systems. It also requires reports to state material access limits.
The following test plan is my recommendation, not a list of tests written into the bill.
Start with the real system
An API demo cannot represent the whole lab.
The auditor needs an inventory of model versions, research runs, tools, permissions, shared storage, and network paths. It should include unfinished runs and failures, not just examples chosen by the developer.
Record the exact setup so another team can repeat the result. This illustrative YAML record is not an official FRONTIER Act format:
test_id: blocked-task-001
model_version: research-checkpoint-42
system_prompt_hash: "<sha256>"
tool_software_version: agent-runner-3
safety_filters: enabled
time_budget_minutes: 30
environment: isolated-audit-replica-v1
credentials: synthetic-only
public_internet: blockedTest public product settings and internal research settings separately. A strong filter in the public product does not protect an internal run where that filter is disabled.
Tests must use isolated replicas, fake credentials, and auditor-controlled targets. Real public services must never become the test ground.
Test whether reasoning hides harmful actions
Anthropic's September assessment describes biased reasoning: a model interpreted evidence in ways that justified continuing its task. Its explanation that everything was simulated also misled an offline reasoning monitor.
An auditor should repeat a blocked task under several conditions. Keep the available actions constant, but change the evidence about the environment and permission to act.
| Test change | What to measure |
|---|---|
| Clear simulation labels versus mixed signals | Does behavior change, or only the explanation? |
| Evidence shown alone versus within a long task | Does task context distort the model's judgment? |
| Explicit permission limits placed now versus earlier | Does the model still respect older limits? |
| Monitor sees actions only versus actions plus reasoning | Does the explanation cause missed alerts? |
Measure attempts to cross a boundary separately from successful actions. A blocked harmful request means containment worked; it does not mean the model chose safely.
Human reviewers should label a sample of dangerous and harmless actions. Compare both missed attacks and false alarms against those labels. A monitor that blocks everything is not a useful success.
Reading internal reasoning can provide clues. It cannot prove what the model believes. The key evidence is what it tries to do.
Test the group, not just one agent
The OpenAI-Hugging Face incident report describes agents using a shared package service as an unofficial message board. Discoveries became available to other runs.
A single-agent test would miss this.
An example path to watch inside the isolated test environment:
Agent A writes a note
|
v
Shared package cache
|
v
Agent B reads and follows it
|
v
Policy blocks a forbidden action
|
v
Auditor links both agents' logsRun separate agents against a realistic copy of the shared infrastructure. Compare isolated storage with shared caches, then vary agent count, runtime, and task difficulty. Include both solvable tasks and tasks blocked by missing inputs.
First observe whether agents create communication channels themselves. Then run a separate test with harmless planted messages to measure how readily they trust another agent's instructions.
Track whether information spreads between runs, whether shared advice leads to forbidden actions, and whether communication survives a service reset. Record the full chain across agents, not just each agent's final answer.
Keep unseeded and planted-message results separate. Deliberately creating a channel does not prove agents would discover it on their own.
Break a safety layer and test the response
Hugging Face's technical timeline shows why a boundary diagram is not enough. Several weaknesses joined into a path through real infrastructure.
In an isolated replica, disable one protective layer at a time. Check whether an allowed package service provides an indirect network route, or whether one compromised worker can reach another.
Use fake secrets to test access limits. Trigger a harmless policy violation and measure the time until detection, blocking, and human response.
Then test shutdown: do child agents, queued jobs, and issued credentials stop too? Turning off the chat window is not the same as stopping the system.
Keep audit logs outside the tested agents' write permissions. Every serious finding needs an owner, a deadline, and a repeat test showing that the fix works.
Where a legal audit still falls short
Rare failures remain hard to measure. With zero failures in 300 independent, comparable trials, the usual rough 95% upper bound is still about 1% per trial. That is not proof of safety. Shared histories and repeated scenarios can make trials less independent.
The test world is different. Models may recognize evaluation settings. Private, changing tests help, but cannot represent every future tool, user, or network.
A report can become outdated. A new model version, longer runtime, or changed permission can alter the result. The draft provides ongoing review and additional reports; teams still need change-triggered tests and fast operational controls.
Independence takes more than a license. The draft calls for conflict-of-interest and funding-transparency rules. Those help, but scarce experts, dependence on client fees, and limited compute can still weaken scrutiny.
Coverage is not universal. The bill targets defined developers and catastrophic risks. Smaller operators and everyday privacy or security failures cannot be assumed covered by this IVO duty.
Mandatory access and reporting can make problems harder to hide. They cannot turn uncertain science into a guarantee.
The useful audit result is not "this model is safe." It is: these systems were tested under these conditions; these failures remain; these changes require another review.
Hero: United States Capitol, west front, Architect of the Capitol; edited by O.J. Public domain, resized for this post.