Technology7 min read

Researchers Used Claude to Hack OpenAI: What Happened

Security researchers used Anthropic's Claude to breach an OpenAI employee's ChatGPT account, exposing private data. Here's what the AI industry hack reveals.

Researchers Used Claude to Hack OpenAI: What Happened

Key takeaways

  1. 1How Researchers Used Claude to Breach OpenAI Systems The mechanics were straightforward by penetration testing standards.
  2. 2OpenAI launched a public bug bounty program in 2023 through Bugcrowd, offering researchers up to $20,000 for critical vulnerabilities.
  3. 3The Irony of a Rival's AI Being Used Against OpenAI The headline tension here is real.
  4. 4The National Institute of Standards and Technology's AI Risk Management Framework, released in 2023, explicitly identifies adversarial attacks and unauthorized system access as priority risk categories for AI deployment.
Sections · 6

A paid security team breached an OpenAI employee's account and read private software information — using their direct competitor's AI tool to do it. The incident, disclosed in September 2026, encapsulates something fundamental about the AI industry's current moment: companies building the most powerful software on earth are still working out how to defend their own infrastructure against those same tools.

How Researchers Used Claude to Breach OpenAI Systems

The mechanics were straightforward by penetration testing standards. A small cybersecurity group gained access to an OpenAI employee's ChatGPT account. Once inside, the researchers could read private software information and propose changes — the kind of access that, in unauthorized hands, could expose proprietary model architectures, internal tooling, or confidential communications.

The route in ran through Anthropic's ecosystem. The research team had been granted access to a specialized Anthropic tool built specifically for security professionals — not a consumer-facing product, but a capability designed to assist with authorized offensive security work. That distinction matters enormously. The fact that Claude hacked OpenAI in this context was not espionage; it was exactly what a well-run vulnerability disclosure program looks like in practice.

Bug bounty programs operate on a principle borrowed from epidemiology: expose the pathogen under controlled conditions before it finds its own way in. The researchers were compensated for their findings, turning a potential exploit into institutional knowledge OpenAI can act on.

Inside the Bug Bounty Program Behind the Hack

Inside the Bug Bounty Program Behind the Hack — a close up of a computer screen with a purple background
Inside the Bug Bounty Program Behind the Hack — a close up of a computer screen with a purple background

Understanding how authorized penetration testing works clarifies what this incident was — and wasn't. Bug bounty programs began scaling seriously in the mid-2010s. HackerOne, one of the largest coordinated disclosure platforms, reported in its annual Hacker-Powered Security Report that organizations now pay researchers hundreds of millions of dollars annually to surface vulnerabilities before malicious actors do. Bounty payouts have tracked closely with the expansion of attack surfaces at large technology companies.

Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026

OpenAI launched a public bug bounty program in 2023 through Bugcrowd, offering researchers up to $20,000 for critical vulnerabilities. The program covers ChatGPT, the API infrastructure, and supporting web properties. Its existence signals that the company understands the basic tradeoff: paying researchers a fraction of what a breach would cost is rational security economics.

The researchers here operated within that framework. They were paid. They had explicit access to the Anthropic tool. Their goal was to surface exploitable weaknesses before bad actors could reach them first. The fact that Claude hacked OpenAI successfully — meaning the attack chain worked — is precisely the point of the exercise. A bug bounty program that never surfaces anything meaningful is either a sign of extraordinary security hygiene or an underpowered program that isn't attracting serious researchers.

What This Reveals About OpenAI's Security Posture

What This Reveals About OpenAI's Security Posture — a computer screen with a quote on it
What This Reveals About OpenAI's Security Posture — a computer screen with a quote on it

That an employee account became a viable entry point is worth examining carefully. Account-level compromises are consistently among the most common attack vectors in the tech industry. Verizon's annual Data Breach Investigations Report has documented for years that credential misuse accounts for a significant plurality of confirmed breaches — recent editions have placed it above 40 percent of all confirmed incidents analyzed across industries.

Employee accounts at AI companies carry unusual privilege. A ChatGPT account belonging to someone on an engineering or policy team may surface internal documentation, model configuration discussions, or early-stage product information valuable to competitors, state actors, or journalists. The attack surface extends beyond technical infrastructure to every authenticated session on every internal tool.

What this incident suggests is not that OpenAI's security is uniquely deficient. It suggests the company operates in the same environment every large technology company does: one where human factors — credential reuse, phishing susceptibility, session management gaps — create persistent exposure regardless of how sophisticated the product being built actually is.

The Irony of a Rival's AI Being Used Against OpenAI

The headline tension here is real. Claude hacked OpenAI — specifically, an Anthropic product served as the instrument through which a breach of an OpenAI system was achieved. The two companies are the most prominent rivals in commercial AI development. They compete for enterprise contracts, developer attention, and talent.

Anthropic built the security-professional tool used in this operation with a stated purpose: enabling authorized offensive security work. That framing reflects Anthropic's public positioning around responsible AI deployment. The company's Responsible Scaling Policy, published publicly, commits to evaluating its models for dangerous capabilities before wider release. Using that same rigor to help identify another company's vulnerabilities is, in a narrow sense, consistent with that mission.

The fact that an Anthropic tool was the vector is not inherently damning for either company. Security professionals routinely use multiple vendors' products in a single engagement. A penetration tester might use one company's reconnaissance tooling, another's exploitation framework, and a third's reporting software. The brand on the tool matters less than whether the operation was authorized, scoped, and properly disclosed.

Still, the optics are striking. In a competitive landscape where both companies regularly publish safety commitments and position their approaches as more responsible than the other's, the image of Claude hacked OpenAI carries a subtext that neither company fully controls once it enters public discourse.

Broader Implications for AI Industry Security

AI companies are high-value targets for reasons that go beyond conventional software security. They hold proprietary model weights representing billions of dollars in compute investment. They store user conversation histories at scale. They manage API keys granting access to generation capabilities that could be redirected for automated fraud, influence operations, or targeted disinformation.

The National Institute of Standards and Technology's AI Risk Management Framework, released in 2023, explicitly identifies adversarial attacks and unauthorized system access as priority risk categories for AI deployment. The framework treats AI security as information security with additional attack surface — not a separate discipline, but an expanded version of familiar problems.

Security researchers who specialize in AI systems have noted publicly that large language models introduce novel social engineering vectors. An LLM with access to internal documents can be prompted in ways that surface information it was not explicitly instructed to protect. The attack surface extends into the model's own behavior, not just the network perimeter surrounding it.

As AI tools become more capable and more embedded in enterprise workflows, the security audit function becomes more critical — and more complex.

What AI Companies and Users Should Take Away

The pragmatic lessons split across two audiences. For AI companies, this incident reinforces that robust bug bounty programs are not optional at scale. Coordinated disclosure — where researchers surface vulnerabilities under confidentiality agreements and receive fair compensation — consistently outperforms purely defensive security postures. The question for OpenAI is whether scope, payouts, and researcher access are calibrated to attract the quality of talent capable of finding meaningful vulnerabilities.

For enterprise users of AI platforms, the incident is a reminder that account security deserves the same attention applied to email and identity infrastructure. Multi-factor authentication, session management policies, and privileged access reviews apply as much to AI platform accounts as to any other high-value system in the stack.

The bigger picture is not alarmist. Claude hacked OpenAI in the most structured, controlled sense of the phrase — researchers did their job, a program worked as designed, and a vulnerability was disclosed before a hostile actor found it first. The industry's challenge is scaling that rigor as the systems themselves scale. At the frontier of AI development, the gap between authorized and unauthorized access is narrow enough that maintaining it requires deliberate, continuous effort — not just from the companies building these systems, but from the researchers willing to test them honestly.


Source: Ars Technica - All content

Published

28 September 2026

Author

Editorial

Comments

No comments yet. Be the first.

Leave a comment