Safety & Security

OpenAI Disrupts Moonshot-Linked Bid to Steal Model Reasoning

OpenAI says it disrupted a coordinated attempt to extract its models' hidden reasoning, tying a core cluster of activity to Moonshot AI after 16,000 requests in two days.

By Rebecca Stone3 min read

Updated

Why it matters

  • OpenAI says it disrupted a campaign that surged to 16,000 requests from over 4,000 users in two days, tracing related activity to a cluster of more than 15,000 users, fully disrupted by July 28.
  • OpenAI attributed a core cluster of the 'adversarial distillation' activity to individuals associated with Moonshot AI, developer of Kimi; Moonshot did not respond to CNBC's requests for comment.
  • OpenAI shared its findings with other AI developers via the Frontier Model Forum and government information-sharing channels; the disclosure follows Anthropic's recent accusation that Moonshot AI and Alibaba used Claude to train their own systems.

OpenAI says it identified and shut down a coordinated campaign to extract the protected reasoning of its AI models, attributing a core cluster of the activity to individuals associated with Chinese startup Moonshot AI, the developer of the Kimi assistant.

The campaign escalated fast. Activity began in early July and surged to 16,000 requests from more than 4,000 users over just two days, according to OpenAI. The company ultimately traced related activity across a cluster of more than 15,000 users and says it fully disrupted the campaign by July 28.

The operators did not breach OpenAI's encryption, databases or stored user conversations, the company said. Instead, they manipulated interactions with its models in an effort to reproduce hidden reasoning in a form visible to the requester. In other words, they tried to get the models to expose the chain-of-thought working that OpenAI deliberately keeps concealed from users.

OpenAI described the activity as "adversarial distillation" — a technique in which one AI model's outputs or reasoning are used to help train or improve another model. The stakes are direct: extracted reasoning could allow a competitor to reproduce advanced capabilities without making the same investment in developing and safeguarding frontier models, which OpenAI said poses potential safety and national security risks.

A pattern of accusations against Chinese AI firms

The disclosure lands weeks after Anthropic, OpenAI's chief U.S. rival, accused several Chinese AI developers — including Moonshot AI and Alibaba — of secretly using its Claude model to help train their own AI systems. Taken together, the two cases mark an escalating effort by leading U.S. labs to publicly frame the extraction of model outputs and reasoning as both a commercial threat and a security issue.

The concern is economic as much as technical. Frontier models cost hundreds of millions of dollars in compute, data and safety work to build. Distillation attacks, if successful, let rivals shortcut that process — reproducing capabilities more quickly and more cheaply than building them from scratch. For U.S. companies competing with a wave of well-funded Chinese startups, that arithmetic shapes everything from API pricing to export-control policy debates in Washington.

OpenAI said it was unclear whether all the operators involved in the campaign were linked to a single actor. But the company attributed a core cluster of the activity to individuals associated with Moonshot AI, the Beijing-founded startup behind the Kimi chatbot and its family of reasoning models.

Moonshot did not immediately respond to CNBC's requests for comment.

Coordinated disclosure through industry channels

OpenAI has not kept the findings to itself. The company said it shared its analysis with other AI developers through the Frontier Model Forum, the industry body created by major labs to coordinate on safety and policy, as well as through government information-sharing channels. That choice signals how seriously the company wants peers and regulators to treat reasoning extraction — not as routine API abuse, but as a category of threat deserving coordinated response.

The timing matters. U.S. policymakers are actively debating how to protect model weights and model outputs from foreign access, and public attributions like this one feed directly into that conversation. An allegation tying a named Chinese company to a distillation campaign gives ammunition to those arguing for tighter controls on model access, usage monitoring and cross-border data flows.

For users of ChatGPT and OpenAI's API, the company's account carries one reassuring detail: no databases were breached and no stored conversations were exposed. The attack targeted the models' hidden reasoning itself, not customer data.

The episode also raises an unresolved technical question for the industry. Chain-of-thought reasoning is hidden partly because exposing it makes models easier to exploit and distill — yet labs increasingly rely on reading that reasoning to monitor models for misbehavior. As distillation attempts grow more brazen, frontier developers will face continued pressure to harden reasoning confidentiality without giving up the visibility their own safety teams depend on.

Original: openai.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

172 articles

Related articles

  1. OpenAI Disrupts Distillation Campaign Tied to Moonshot AI
  2. OpenAI Bans Accounts Reviving Russia's 'Stop News' Influence Operation
  3. OpenAI models broke out of isolation and breached Hugging Face
  4. OpenAI Halts Training of Its Most Powerful Models
  5. AI Models Keep Cheating on Tests, and Researchers Are Quitting

« Previous article