Security researchers used Claude to help them hack into OpenAI
A team of three security researchers successfully hacked into OpenAI employee accounts and accessed its GitHub repository, "Monorepo," with the assistance of Anthropic's Claude Opus 4.8 and 5.
Intelligence analysis by Gemini 2.5 Flash

The researchers, operating as Hacktron, exploited a vulnerability in Discourse, a third-party service hosting OpenAI's community forums, specifically targeting how it processed HEIF image files. They achieved remote code execution (RCE) and gained access to OpenAI's instance, demonstrating the breach by sending a pull request from an employee's Codex account.
Imagine a super-smart computer helper, like a really clever robot, that helped some tech detectives find a secret back door into a big company's online clubhouse, OpenAI. They tricked a security robot at the clubhouse's front gate by showing it a special, tricky picture. Once inside, they could peek into the company's secret recipe book for making its smart computer brains, showing everyone that even big companies need to check their locks carefully.
Analysis
The recent breach of OpenAI's systems by a small team of security researchers, Hacktron, underscores a critical juncture in cybersecurity, where artificial intelligence models are not merely targets but active participants in the attack chain. The incident, which saw researchers gain access to OpenAI's GitHub repository, "Monorepo," containing what are described as "algorithmic secrets," was achieved with remarkable speed and efficiency, largely attributed to the assistance of Anthropic's Claude Opus models.
Claude Opus 5
The pivotal role of Anthropic's Claude Opus 4.8 and 5 in this security breach cannot be overstated. According to Hacktron, the team leveraged Claude Opus 5, which launched on July 24th, to achieve remote code execution (RCE) on Discourse Cloud and access OpenAI's instance by 10 AM the very next day. This rapid turnaround suggests that advanced large language models (LLMs) are becoming increasingly sophisticated tools for identifying and exploiting vulnerabilities, significantly accelerating the reconnaissance and exploitation phases of cyberattacks.
The researchers spent less than $3,000 in tokens, indicating that the computational cost of using these powerful AI assistants for complex security tasks is relatively low. This accessibility, combined with the models' ability to quickly adapt exploits, presents a dual-edged sword: while it empowers ethical hackers to uncover critical flaws, it also lowers the barrier for less scrupulous actors to develop and deploy sophisticated attacks, potentially at scale.
HEIF Heist
The technical vector for the breach was dubbed the "HEIF Heist," targeting a specific vulnerability within Discourse, the third-party forum software used by OpenAI. The exploit capitalized on an issue with how Discourse processed HEIF (High Efficiency Image File Format) images. By crafting a corrupted image file, the researchers were able to trigger a remote code execution, effectively taking control of the server hosting OpenAI's community forums.
This method highlights the often-overlooked attack surface presented by third-party services and seemingly innocuous file processing functionalities. The ease with which this exploit could be adapted to other major companies, including Slack, Meta, and GitHub Ent, within "only one or two days," demonstrates a systemic vulnerability that extends beyond OpenAI. The fact that only one target, Shopify, reportedly detected the exploit further emphasizes the stealth and effectiveness of this AI-assisted attack methodology.
Monorepo
The ultimate objective and significant outcome of the Hacktron team's efforts was gaining access to OpenAI's GitHub repository, known as "Monorepo." This repository is reportedly a treasure trove of "OpenAI’s algorithmic secrets," implying it contains proprietary code, models, and potentially sensitive research data. While the researchers stopped short of directly accessing internal code, they proved their access by submitting a pull request from an employee's Codex account.
Access to such a critical repository could have catastrophic implications if exploited by malicious entities, ranging from intellectual property theft to the introduction of backdoors or the manipulation of AI models. The incident serves as a stark reminder that even leading AI companies, with their inherent focus on advanced technology, are susceptible to fundamental cybersecurity flaws, especially those residing in their broader IT infrastructure and third-party integrations. The $6,500 bug bounty paid by OpenAI for the disclosure, while a positive outcome for responsible disclosure, pales in comparison to the potential value of the compromised data.
Key points
- Three security researchers from Hacktron used Anthropic's Claude Opus 4.8 and 5 to hack into OpenAI employee accounts.
- The breach was achieved by exploiting a vulnerability in Discourse, a third-party service hosting OpenAI's community forums, related to HEIF image processing.
- Researchers gained access to OpenAI's GitHub repository, "Monorepo," which reportedly contains "OpenAI’s algorithmic secrets."
- The "HEIF Heist" project took less than 72 hours to execute and cost under $3,000 in AI tokens.
- The vulnerabilities were reported to Discourse and OpenAI and have since been fixed, with OpenAI paying Hacktron $6,500 for the bug.
The successful ethical hack by Hacktron led to the prompt identification and patching of critical vulnerabilities in Discourse and OpenAI's systems, demonstrating the value of responsible disclosure in enhancing cybersecurity. OpenAI's payment of a bug bounty further encourages security researchers to find and report flaws, ultimately making AI infrastructure more secure.
The ease and speed with which a small team, using readily available AI tools, could achieve remote code execution and access sensitive data highlights a significant and growing threat. This incident suggests that malicious actors could leverage similar AI capabilities to launch sophisticated attacks, potentially compromising critical AI models and data with minimal effort and cost.



