
Claude Helped Hack OpenAI: What the Hacktron Breach Changes
Yoni Fraimorice
A three-person security team used Claude Opus 5 to help break into OpenAI.
This was authorized security research. Hacktron AI reported the weaknesses, stopped after proving the impact, and received a $6,500 bounty. No customer data or model weights were reported stolen.
Still, the incident matters. The team moved from an image upload to remote code execution, employee ChatGPT and Codex account access, and an internal repository proof in less than 72 hours.
The software bugs were serious. The larger warning is economic: frontier AI models are making difficult exploit development faster and cheaper.
The chain started with an image
OpenAI uses Discourse for its community forum. Most uploaded images went through FastImage checks. HEIC and HEIF files followed another route because FastImage did not support them: Discourse passed them to ImageMagick, which used the native libheif library.
That created this path:
Attacker-controlled HEIC file
-> Discourse upload
-> ImageMagick
-> vulnerable libheif decoder
-> memory corruption
-> code execution on the forum serverHacktron found a heap buffer overflow that provided out-of-bounds memory read and write capabilities. The relevant upstream code had already changed, but the commit was not clearly marked as a security fix and did not have a CVE at the time. Standard security-backport processes therefore missed it.
The Discourse Docker image used Debian 12 with vulnerable libheif 1.19.7. Even Debian 13 still carried an affected version. A dependency can be unsafe while every CVE-based scanner reports green.
Opus 5 turned a bug into a working exploit
Hacktron says Claude Opus 4.8 found the missing security backport and produced an exploit when address space layout randomization, or ASLR, was disabled. Several attempts to make it reliable with ASLR enabled failed.
Anthropic released Claude Opus 5 during the investigation. The researchers gave the new model the same problem.
According to Hacktron's technical disclosure, Opus 5 produced a working ARM64 exploit for a local Mac in about three hours. It then helped port the exploit to the x86-64 and jemalloc environment used by Discourse.
By 6:00 a.m. on July 25, the team had local code execution through an image upload. Opus initially refused to attack a remote system, so the researchers presented their own Discourse Cloud instance as a CTF target. The agent achieved code execution and read /etc/hosts as proof.
The team then used the generated script against OpenAI's forum and obtained remote code execution and administrative access.
This was not fully autonomous hacking. Skilled researchers selected the target, handled safety decisions, and connected the stages. The model still performed work that normally requires rare memory-exploitation expertise.
The identity layer multiplied the damage
Compromising a forum should not automatically compromise an employee's development tools. Hacktron found a separate OpenAI single sign-on issue that, according to the researchers, allowed takeover of ChatGPT and Codex accounts belonging to active forum members without another user interaction.
The exact SSO flaw was not publicly disclosed, so claims about its precise cause would be speculation.
One affected employee account had Codex connected to OpenAI's GitHub organization. To prove access without reading internal source code, the researchers instructed Codex to open one harmless pull request in OpenAI's private monorepo. They then stopped testing and updated their report.
one weak application
-> federated identity
-> agent account
-> connected services
-> internal code and dataAn AI agent with many connectors is a high-value identity with inherited reach.
What this proves about attack cost
Hacktron reports that the OpenAI chain took less than 72 hours, with only a few hours of direct human work. Its wider HEIF Heist project involved three researchers, lasted about two months, and cost less than $3,000 in model tokens. Adapting the exploit to another target often took one or two days.
One case cannot prove that every advanced exploit is now easy. The comparison between Opus 4.8 and Opus 5 was not a controlled benchmark. Human expertise remained essential.
But the direction is clear. A public memory bug could remain difficult to weaponize because few people had the time and skills to build a stable, target-specific exploit.
Frontier models turn part of that scarce expertise into compute. Defenders should assume that complex bugs will be operationalized faster, by smaller teams, against more targets.
Defense 1: scope agent permissions
Do not give a general-purpose agent every permission its human owner has.
- Use a separate service identity for each agent and environment.
- Grant repository, email, chat, and cloud access independently.
- Require human approval for new connectors and sensitive write actions.
- Use short-lived credentials and remove them when a task ends.
- Block network destinations and tools that the current task does not need.
- Keep agent audit logs outside the agent's write permissions.
A useful rule is: compromising one agent session should not compromise the employee's whole digital identity.
Defense 2: clean up federated identity
Every SSO trust relationship expands the possible blast radius.
Inventory every application that accepts your identity provider and every service reachable through connected agents. Validate token issuer, audience, expiry, signature, and intended client. Do not accept a token issued for one application in another.
Use step-up authentication before opening internal repositories, sending email, changing cloud resources, or adding connectors. Revoke downstream sessions when the upstream account is disabled. Alert when a low-risk application suddenly reaches a high-value service.
Most importantly, test the full path. A secure identity provider cannot save an application that interprets its tokens incorrectly.
Defense 3: shrink image-parsing exposure
Image upload is native-code execution applied to attacker-controlled bytes.
- Accept only formats the product actually needs.
- Disable HEIF and AVIF decoding if users do not require them.
- Keep image workers isolated from secrets, internal networks, and control planes.
- Use ephemeral sandboxes with CPU, memory, file, and time limits.
- Re-encode uploads into a simple internal format before further processing.
- Track upstream commits and distribution advisories, not only CVEs.
- Monitor decoder crashes and repeated malformed uploads as security signals.
Discourse's advisory tells self-hosted operators to rebuild with the latest Docker image. Discourse also added sandboxing around image processing as defense in depth.
The new security baseline
The Hacktron incident does not mean AI can replace an expert offensive-security team. It means a small expert team can now attempt far more work.
That changes reasonable defense assumptions. Patch windows must shrink. Agent identities must have narrow permissions. SSO boundaries must be tested as attack paths. Complex file parsers must run as if compromise is expected.
The model was not the vulnerability. It was the force multiplier that made several old security failures cheaper to connect.
Sources
- Hacktron AI: Hacking OpenAI
- Discourse advisory: RCE via malformed HEIF file
- Debian DSA-6417-1: libheif security update
- Anthropic: Introducing Claude Opus 5
- VentureBeat: OpenAI hacked by researchers using Claude Opus 5
- The Wall Street Journal: Hackers used Claude to break into OpenAI
Hero image: Cyber Shield defensive exercise, U.S. Army photo by Sgt. 1st Class Jon Soucy, public domain.