Safety & Security

Security Researchers Used Anthropic's Claude to Hack Into OpenAI

A paid vulnerability program let a small security team use Anthropic's tool to access an OpenAI employee's ChatGPT account and read private software data.

Researchers used Claude to hack OpenAI
Researchers used Claude to hack OpenAIAI-generated
By Sophie Lindqvist4 min read

Updated

Why it matters

  • A small cybersecurity group accessed an OpenAI employee's ChatGPT account using an Anthropic security tool.
  • The researchers could read private software information and suggest changes.
  • Anthropic paid the researchers through a program to find vulnerabilities before malicious actors could exploit them.

A small cybersecurity group broke into OpenAI using software built by Anthropic, OpenAI's chief rival, exposing weaknesses in the ChatGPT maker's defenses at a moment when the industry's safety practices face intensifying scrutiny.

The intrusion gave the researchers access to an OpenAI employee's ChatGPT account. From that vantage point, they could read private software information and suggest changes, according to details of the incident.

The breach was authorized. The researchers had been granted access to an Anthropic tool built specifically for security professionals, and Anthropic paid them for the work. The operation ran under a program designed to surface vulnerabilities before malicious actors could exploit them — a bug-hunting arrangement, not an attack.

The episode matters beyond the irony of one AI lab's model being used to test another's defenses. OpenAI and Anthropic are locked in direct competition in the market for frontier AI systems, and both are spending heavily on safety infrastructure. That a small outside team could reach an employee account with the ability to view internal software information and propose modifications illustrates how far ordinary workplace tools — in this case ChatGPT itself — have become part of a company's attackable surface.

The fact that the researchers used Anthropic's tooling adds a competitive dimension. Anthropic positioned its security-focused offering as a product for exactly this kind of professional work, and the results demonstrate that such tools are capable of driving real penetration efforts against a major AI company's internal systems. For enterprises weighing whether AI-assisted security testing belongs in their programs, the OpenAI case is a concrete data point.

The stakes are heightened by the political environment. Leading AI companies face mounting scrutiny over safety from regulators and lawmakers, and incidents like this one feed directly into that debate. A company that cannot keep its own employee accounts secure will struggle to argue that its broader safety commitments are airtight. Conversely, the outcome here also shows the value of paid vulnerability programs: the weakness was found by researchers compensated to look for it, rather than by an attacker motivated to exploit it.

The mechanics of the finding are straightforward. The security group obtained access to the ChatGPT account of an OpenAI employee. ChatGPT accounts, like other workplace software, can hold traces of private company information — in this case, private software details the researchers were then able to read. The researchers went a step further and suggested changes, demonstrating a level of access that extended beyond passive observation.

Anthropic's role was limited to supplying and paying for the capability. The company's tool was designed for security professionals, and the researchers operated it within the terms of a vulnerability-finding program. There is no indication in the reported details that Anthropic targeted OpenAI directly; rather, its product was the instrument the researchers chose or were licensed to use.

For OpenAI, the incident is a reminder that the security perimeter now includes the AI assistants its own employees use daily. As companies across the economy fold chatbots and AI agents into routine workflows, the accounts tied to those tools become repositories of institutional knowledge and, potentially, vectors for manipulating internal processes. An account that lets a user suggest changes to software is not merely a leak risk — it is a potential supply-chain risk.

The industry-wide implication cuts both ways. On one hand, an AI model proved capable enough to support a successful intrusion into a top AI lab's systems, which will intensify concerns that the same capabilities are available to less well-intentioned actors. On the other, the discovery happened inside a structured, paid program — the mechanism working as intended, finding the flaw before criminals did.

That dual reading is likely to shape how both companies and regulators respond. Expect greater pressure on AI firms to run similarly aggressive internal and third-party testing of their own deployments, and expect Anthropic to cite this result as evidence that purpose-built security tools with professional guardrails can deliver findings that traditional audits miss. For OpenAI, the immediate task is remediation: closing whatever path let an employee account expose private software information to outsiders, even authorized ones.

The larger question the incident leaves open is whether the pace of AI-assisted security testing can keep up with the pace of AI-assisted attack. This time, the defenders paid first and found the hole. The next team to reach an account like this may not send a report.

Source: Ars Technica AI

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. OpenAI adds Lockdown Mode and risk labels to ChatGPT
  2. OpenAI agents leaked 53 ChatGPT user images
  3. OpenAI models broke out of isolation and breached Hugging Face
  4. OpenAI Launches Advanced Account Security for High-Risk ChatGPT Users

Next article »