Claude Helped Hack OpenAI in 72 Hours. Is Expertise Becoming Optional?

A three-person security team used frontier AI to turn a malformed image upload into a route toward OpenAI’s internal systems in less than 72 hours—but the humans still drove every consequential move.

A malformed image became the first step toward OpenAI’s internal systems. Hacktron says its three-person security team linked a HEIF image-processing flaw to a weakness in OpenAI’s single sign-on, reaching multiple employee ChatGPT accounts and an internal repository in less than 72 hours. That chronology comes from the researchers’ disclosure; the public record does not include OpenAI’s internal logs. Hacktron AI — Hacking OpenAI The Verge — Security researchers used Claude to help hack OpenAI

The team says it stopped after proving access by directing an employee’s Codex account to open a harmless pull request in an internal monorepo. It says it did not read internal code. OpenAI later confirmed that the reported vulnerabilities had been addressed, according to The Guardian. Hacktron AI — Hacking OpenAI The Guardian — OpenAI ethically hacked with help of Claude

The image was the doorway

The chain began with malformed HEIF uploads. Discourse’s security advisory confirms that CVE-2026-32882 allowed remote code execution through those files, rates the flaw 8.8 on the CVSS scale and says supported versions added sandboxing for image processing. In ordinary terms, a feature meant to display a picture could be made to run unwanted instructions. Discourse security advisory GHSA-vhm9-85gw-x335

Hacktron says that foothold was connected to an OpenAI single-sign-on weakness, allowing the researchers to reach employee accounts. The disclosure does not show that sensitive source code was copied or that the operation was autonomous. It describes a controlled proof of access instead. Hacktron AI — Hacking OpenAI

AI shortened the hard part

According to Hacktron, Claude Opus 4.8 helped identify missing security backports but struggled to make the exploit reliable against ASLR, a protection that randomizes where software is placed in memory. Claude Opus 5 later produced working local and Discourse Cloud exploits. Human researchers still selected the target, supplied direction, judged failures and connected the separate weaknesses. Hacktron AI — Hacking OpenAI

The model record is not uniform across reports: The Guardian says the team later relied largely on OpenAI’s GPT-5.6 Sol as well as Claude. The evidence therefore supports AI-assisted exploit development, not a clean story in which one system independently hacked its rival. The Guardian — OpenAI ethically hacked with help of Claude Hacktron AI — Hacking OpenAI

The durable change is speed

The practical consequence is a shorter distance between a vulnerability and a working demonstration. In this case, models helped researchers search for missing fixes, generate code and iterate on failures. The operation still depended on skilled judgment and a carefully chosen chain, so the evidence does not show that equivalent attacks are now available to anyone. Hacktron AI — Hacking OpenAI

Hacktron says OpenAI fixed its side roughly 14 hours after the initial report and later paid $6,500 for the finding. The episode points less to machines replacing hackers than to a changed balance of effort: experienced researchers remain essential, while AI gives their decisions more reach and leaves defenders less time to respond. Hacktron AI — Hacking OpenAI The Guardian — OpenAI ethically hacked with help of Claude

Avatar photo
estebanro_admin
Articles: 1

Newsletter Updates

Enter your email address below and subscribe to our newsletter