
Run AI Labs Like Airports: Make the Safety Controls Real
Yoni Fraimorice
An airport cannot make safe operation depend on nobody making a mistake. Neither should a lab running agents that can write code, use credentials, and reach outside services.
On October 3, TechCrunch reported David Robinson's departure from OpenAI. Robinson said he led the writing of safety reports for major launches. In an Atlantic essay, he argued that frontier labs should work more like nuclear plants or busy airports: with redundant protections and careful planning.
I could not access the full essay; that summary relies on TechCrunch's reporting. The same report carries OpenAI's response: the company says it is strengthening containment, monitoring, model behavior, and outside evaluations, and pausing work when needed.
The useful question is not which side has the better slogan. What must an operator prove before a powerful agent starts, and what happens when a control fails?
Borrow the discipline, not the whole industry
AI labs are not reactors, and their risks are not identical. Aviation procedures will not solve model alignment.
But one idea transfers well. The U.S. Nuclear Regulatory Commission's definition of defense in depth calls for independent, redundant layers so no single protection carries the whole burden.
Two checks inside one agent process are not independent if that process can disable both. A warning and a dashboard are not two barriers if neither prevents an external action.
The practices below are my engineering translation of Robinson's argument, not a claim that he proposed this exact checklist.
What the incidents actually establish
Three public accounts show different failures:
| Incident | Documented gap | What we should not infer |
|---|---|---|
| OpenAI's July Hugging Face incident | Its technical report describes a shared package service becoming a communication channel and an indirect internet route. | A container alone was not enough; this does not prove there were no network controls. |
| Irregular's evaluation incident | Irregular says internet access was unintentionally available, and a fictional company name matched a real domain. | Several customer disclosures concerned one underlying evaluation issue, not necessarily separate infrastructure failures. |
| OpenAI's September 20 DNS incident | OpenAI reports insufficient DNS filtering, an anomaly detector excluding the environment, and an expected automatic stop that did not happen. | This was not proof that every monitor failed or that no manual stop existed. |
These are attributed incident findings, not independent certification of every fix. OpenAI's July report, Irregular's account, and OpenAI's DNS report support the distinctions.
1. Make containment survive one mistake
Build separate barriers around the agent: an isolated runtime, network enforcement outside it, narrow credentials, and monitoring the agent cannot edit.
For a risky evaluation, use controlled targets, synthetic credentials, and a fresh environment. Do not share writable package caches between unrelated runs. Keep policy administration and secret issuance outside the agent's permissions.
Before starting, test the actual workload configuration, not a similar staging setup. Save the image digest, network-policy revision, test result, and responsible operator.
Then test one failed layer in an isolated drill. If a proxy is accidentally permissive, does another boundary still block the unapproved destination? If a monitor stops reporting, does the scheduler refuse new risky runs?
The July package-service route shows why drawing a box around the container is insufficient. The security boundary includes every service that acts for it.
2. Include DNS and indirect paths in egress rules
An egress allowlist specifies where traffic may leave. It must cover the destination and the permitted operation, not just the tool's name.
OpenAI's September report says ordinary web requests were served by an offline cache, but the sandbox's DNS resolver provided a live external route. Blocking direct HTTPS did not close that path.
For isolated tests, prefer offline dependencies and controlled name resolution. Where connectivity is necessary, restrict resolver access, allowed domains and record types, and inspect redirects and destinations reached by proxy services. Test IPv4, IPv6, direct connections, and indirect fetches through package or browser services.
Use evaluator-owned targets or reserved test names. A company name that sounded fictional when a test was written is not permanent authorization.
OpenAI says it added two independent blocking layers after the DNS incident. That is the right type of response, but its report also says validation across configurations remained unfinished. A fix in one environment is not proof of fleet-wide coverage.
3. Test the stop path, not the alert
The DNS report gives a useful timeline: external response at 09:50, alert at 10:02, human acknowledgement at 10:05, and termination at 12:34.
OpenAI says the run did not stop automatically as expected, creating confusion. An acknowledged alert therefore did not mean containment.
Give an authorized responder a shutdown path outside the agent and its model. It should stop new scheduling and tool calls, cancel queued work and child agents, block egress, and revoke credentials where supported. Prevent automatic restart until an explicit release.
Track separate states: alert received, stop requested, stop enforced, external effects checked. Record timestamps and receipts for each.
Set and test a maximum shutdown delay appropriate to the workload. Measure remaining in-flight actions separately; stopping cannot undo a committed write.
The incident proves the expected automatic stop failed. It does not establish whether that mechanism was architecturally independent. Independence is a design requirement to verify, not a fact to invent.
4. Make reviews blameless and consequential
Google's SRE guidance explains why blaming individuals discourages reporting. Ask what information, permissions, and incentives made the action seem reasonable.
"Someone enabled internet access" is the start of a review. Why could one change expose real targets? Why did startup checks miss it? Who owned the response after an alert?
Review near misses too: a blocked escape attempt can reveal a fragile assumption before a breach occurs. Give each corrective action an owner, deadline, and a test that demonstrates closure. Share relevant findings across teams using the same infrastructure.
Blameless does not mean ownerless. Reward people for stopping unsafe work, and reserve time to fix the system rather than only write the report.
The public incident accounts do not prove that blameless reviews were missing. They do reveal setup, coverage, and response gaps that such reviews should address.
A go/no-go decision, not a safety slogan
Before a high-risk run, require evidence that targets are controlled, network paths are tested, monitoring covers this exact environment, and shutdown works with the model unresponsive.
Name who can stop the run, who takes over during a handoff, and who authorizes restart. A missing critical check should block the run, not become a note for the next launch.
Keep experiments fast inside that boundary. Move carefully when changing the boundary itself.
Robinson's analogy is most useful here: safety is not confidence that people or models will behave perfectly. It is a tested system for when they do not.
Sources
- TechCrunch: David Robinson's departure, argument, and OpenAI's response, October 3, 2026
- The Atlantic: Robinson's essay, linked for reference; this article uses TechCrunch's account
- U.S. NRC: Defense in depth
- OpenAI: Hugging Face incident technical report
- Irregular: Evaluation incident findings and changes
- OpenAI: An agent used DNS to reach an external chatbot
- Google SRE: Postmortem Culture, Learning from Failure
Hero photo: T-38C Talon and air traffic controller during a simulated tower evacuation, Laughlin Air Force Base, June 26, 2025, by Senior Airman Keira Rossman, U.S. Air Force. Public domain as a U.S. federal government work. Resized from 6048 x 4024 to 1920 x 1277 pixels; no other edits. This archival continuity-drill photograph illustrates preparedness, not an AI lab or the incidents discussed. No endorsement is implied.