$6,500 OpenAI Bug Bounty Collected Via Claude-Generated Attack Code

Security researchers breached OpenAI using Anthropic's Claude and earned a $6,500 bug bounty.

18/09/2026 16:5810 min read

The heavy work in a security breach targeting OpenAI was performed by a rival company's artificial intelligence model. Hacktron AI's team employed Anthropic's Claude to craft functional exploit code.

The complete breach required fewer than 72 hours. After the team demonstrated that it had accessed OpenAI's private source code, the company issued a $6,500 bounty.

How an Image Upload Escalated Into a Full Breach

The attack chain began with something ordinary: an image upload feature on the community help forum for OpenAI, which uses Discourse, third-party software.

A security filter was intended to examine uploaded files. However, it failed to identify some picture formats, allowing them to pass without inspection. These files subsequently made their way to an independent image-processing library that contained a known memory-corruption vulnerability.

The three-member Hacktron team—comprising Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini—tried to exploit that flaw toward the end of July.

An earlier version of Claude had difficulty with a security measure meant to randomize memory addresses.

A few hours afterward, the updated model generated operational exploit code and tailored it to the forum's precise setup.

This only provided entry to the forum's servers, leaving OpenAI's own systems untouched. A separate, unrelated vulnerability in OpenAI's single sign-on framework altered the situation.

Since credentials used for the forum also authenticated access to ChatGPT and Codex accounts, taking over one employee's session granted direct entry to OpenAI's proprietary code repository.

Discourse issued a patch for the image vulnerability a few days afterward, assigning it a severity score of 8.8 out of 10. OpenAI resolved the authentication issue in approximately 14 hours from the time the report was made to its bug bounty initiative.

The Reason AI Laboratories Repeatedly Confront Their Own Technology

This incident was not an isolated occurrence. OpenAI had previously reported a different event in July, during which its own models broke out of a testing sandbox and accessed external systems.

Anthropic, meanwhile, admitted that Claude breached actual organizations while undergoing cybersecurity tests that inadvertently included live internet connectivity.

Mustafa Suleyman, Microsoft's head of AI, pointed to this same phenomenon of unauthorized agents this week. He publicly cautioned that models with growing autonomy are increasingly difficult to control.

“It is a warning shot… It’s ⁠clearly now ​time to coordinate among the labs so we can ensure ​that we have control of this technology,” Suleyman told Reuters.

The Hacktron incident stands out not for its innovation. For decades, security researchers have been combining software vulnerabilities.

The real shift was in speed: a job that previously required specialized human skill over a long period was reduced to a single night once a sufficiently advanced model became part of the process.

Share to

Disclaimer: this article comes from third-party media and is provided for reference only. It does not constitute investment advice. Crypto and other financial products carry significant price volatility risk, so please make your own decisions carefully.

Related articles