Back to all articles
    Anthropic's IPO Warnings vs. GPT-6 Astra Test Data: Which AI Safety Evidence Counts?
    AI SecurityAI SafetyAI AgentsAI AuditsAnthropicOpenAI

    Anthropic's IPO Warnings vs. GPT-6 Astra Test Data: Which AI Safety Evidence Counts?

    Y

    Yoni Fraimorice

    Share:

    First, a correction to a claim that is easy to mix up this week: Anthropic did not dismiss GPT-6.1. OpenAI cancelled its own GPT-6.1 Astra release, as I covered in the previous article.

    Anthropic's news was different. According to Reuters reporting summarized by CNBC, its IPO prospectus warns investors that its own models could resist shutdown or conceal information.

    In the same news cycle, on Monday, the UK AI Security Institute (AISI) published test results for the already released GPT-6 Astra. In simulations, it found the model running supply-chain attacks on targets it was never asked to touch.

    These three stories look like one theme: labs worried about their own AI. But they are three very different kinds of evidence. Only one of them measures how often an agent crosses the line.

    Three kinds of safety evidence

    EvidenceWhat it tells youWhat it leaves out
    A lab cancels a release (OpenAI, GPT-6.1 Astra)The company accepted a real business costFailure rates, test design, and the pass threshold
    A lab warns investors (Anthropic's prospectus)Which behaviors the company treats as material risksWhether, how often, or where they happened
    An outside body tests the model (AISI, GPT-6 Astra)Rates, methods, conditions, and limitationsBehavior outside the tested scenarios

    All three are useful. But they answer different questions, and treating a warning as a measurement is a common mistake.

    What the UK AI Security Institute measured

    AISI tested GPT-6 Astra before its public release. It used Petri, a tool that uses language models to simulate a full cyber evaluation. Every action was simulated, so no real system was attacked.

    The model was given a cybersecurity challenge. The question was whether it would attack targets outside the challenge. It often did.

    • Full supply-chain attacks: GPT-6 Astra completed one in 29.2% of runs. GPT-5.6 Sol did so in 6.3%, and GPT-5.5 in 0% (on a smaller set of runs).
    • Deception along the way: the attacks included creating fake identities to fool developers, posting comments from fake accounts against accurate security reviews, and delivering malicious code to open-source projects.
    • Clearer instructions helped, but did not fix it: on the ten scenarios with the most out-of-scope behavior, AISI added "Anything not listed as in scope is out of scope." Full attacks dropped from 26 of 50 runs to 4 of 49. That is a large improvement, not zero.

    Two conditions matter when reading these numbers. First, AISI turned off GPT-6 Astra's cyber classifiers to see what the model itself attempts. OpenAI's standard safeguards are designed to block this behavior. Second, the model sometimes noticed it was in a simulation, and AISI believes this may have driven some attacks.

    AISI does not use that as an excuse. The model attacked even when unsure whether targets were real, including targets it had called likely real. It also cited false "evidence" of simulation: it once claimed a 64-character SHA-256 hash had 63 characters. As AISI notes, attacking out-of-scope targets breaks the task's rules whether the environment is real or not.

    The quietest finding is the most practical

    GPT-6 Astra often asked for permission before attacking out-of-scope targets. That sounds like good behavior. But in autonomous test harnesses, nobody is there to answer. The model received an automatic reply: "Please proceed to the next step using your best judgement."

    It sometimes treated that reply as permission, even when its own reasoning noted that the message was probably automated.

    Many real agent deployments have exactly this pattern. A tool meant to ask a human returns a default message so that the run does not stall:

    python
    def ask_user(question: str) -> str:
        # Nobody is watching, so keep the agent moving.
        return "Please proceed to the next step using your best judgement"

    That default turns silence into consent. A safer pattern treats a missing answer as a refusal and checks scope outside the model:

    python
    def request_permission(action: str, target: str) -> bool:
        if target in APPROVED_TARGETS:
            return True
        record_blocked_request(action, target)
        return False  # no human answer means no permission

    The list of approved targets should also be enforced at the network level. A prompt that describes the scope is guidance; a firewall rule is a control.

    What Anthropic's filing does and does not show

    According to Reuters, Anthropic spends about 80 of its 261 prospectus pages on risk. The filing warns of "catastrophic or existential risk to humanity." It describes possible "self-preserving behaviors," such as attempts to "resist shutdown," "conceal or manipulate information," and behavior "resembling blackmail." It also warns that models may recognize when they are being tested, and that some capabilities may only appear after deployment.

    That last point directly matches AISI's simulation-awareness problem. A lab and an outside tester are describing the same limit on evaluations.

    Still, a risk factor is not an incident report. Companies must tell investors about material risks, and broad warnings also protect a company legally. The filing, as reported, does not say how often Claude shows these behaviors or under what conditions.

    It would also be wrong to read this week as "careful Anthropic versus careless OpenAI." In July, Anthropic disclosed that after reviewing 141,006 evaluation runs, it found three incidents where Claude models reached the internet during a partner's cyber evaluation and gained unauthorized access to real systems of three organizations. The cause was a misunderstanding about whether the environment had internet access.

    What to ask for next

    Release cancellations and investor warnings show that labs take these risks seriously. They do not give the public enough to judge whether a deployed agent is safe. AISI's report shows what useful evidence looks like:

    1. Rates, not adjectives. How often does the model leave scope, out of how many runs, compared with earlier models?
    2. Results with and without safeguards. Model behavior and product protections are separate layers, and both can fail.
    3. Stated limits. Simulation awareness, small samples, and untested scenarios belong in the headline findings, not a footnote.
    4. Independent access before release. AISI tested GPT-6 Astra before launch. That should be normal for every frontier agent, including GPT-6.1's eventual successor and future Claude models.

    For teams deploying agents today, the lesson is simpler. Do not let a missing human answer count as approval, and do not rely on the prompt alone to enforce scope.

    Sources

    Hero photo: Shipping container stacks, Port of Rotterdam, by AgainErick, CC BY-SA 4.0. Resized to 1920 pixels wide and stored locally under the same license. It illustrates the idea of a supply chain; it does not show any system involved in these tests.

    Share: