On September 18, 2026, The Guardian reported that security startup Hacktron AI used Anthropic’s Claude and OpenAI’s GPT-5.6 Sol to compromise multiple OpenAI employee ChatGPT accounts and reach internal code repositories under a bug-bounty program. Business Today separately detailed how the researchers accessed an employee’s ChatGPT account linked to GitHub, confirming OpenAI paid a 6,500 dollar bounty and says it has fixed the vulnerabilities.
This article aggregates reporting from 3 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
This episode is a sharp reminder that frontier AI models are not just targets for attackers but powerful tools for offense. Hacktron’s team used Anthropic’s Claude and OpenAI’s own GPT-5.6 Sol to pivot from a staff forum to internal ChatGPT accounts and source code, compressing what used to be a months-long red‑teaming effort into days. That is exactly the kind of real-world capability jump security researchers have been predicting once agentic models can reliably write and reason about complex exploit chains.
For the race to AGI, the hack underscores a paradox: as labs build more capable systems, they also hand adversaries smarter tools for discovering and weaponizing their own weaknesses. OpenAI’s bug-bounty framing is important, but the optics are bad: a rival’s model helped expose deep flaws in OpenAI’s access controls. Strategically, this strengthens the argument that agent security and identity layers have to evolve as fast as model capabilities. It also shows that independent red‑team boutiques equipped with commodity frontier models can now pressure even top labs, shifting some power away from internal safety teams and toward a broader security ecosystem.

