Safety & Security

OpenAI Disrupts Distillation Campaign Tied to Moonshot AI

OpenAI disrupted a distillation campaign that extracted protected reasoning via 16,000 requests from 4,000+ users, attributing a core cluster to individuals tied to Moonshot AI.

Disrupting a coordinated model-distillation campaign
Disrupting a coordinated model-distillation campaignAI-generated
By Sophie Lindqvist5 min read

Updated

Why it matters

  • OpenAI attributes a core cluster of the distillation campaign to individuals associated with Moonshot AI, developer of Kimi.
  • Activity began July 1; spikes on July 24–25 involved 16,000 extraction-pattern requests from over 4,000 users; a related cluster of 15,000+ users was disrupted by July 28.
  • Operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it; no encryption was broken and no database was compromised.

OpenAI says it disrupted a coordinated campaign that extracted protected chain-of-thought reasoning from its models, and it attributes a core cluster of the activity to individuals associated with Moonshot AI, the Chinese developer of the Kimi model family. The company disclosed the operation in a post titled "Disrupting a coordinated model-distillation campaign," marking one of the most specific public attribution statements a frontier lab has made about adversarial distillation to date.

The campaign's activity began on July 1 at low volume. OpenAI then observed high-volume spikes on July 24 and 25, consisting of 16,000 requests using a relevant extraction pattern from more than 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which OpenAI says it fully disrupted by July 28. The earliest observed activity occurred in the first week of July.

Adversarial distillation is the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model's internal record for working through a task. Extracting it can reveal information withheld from the final answer and help others reproduce the model's capabilities.

The stakes are straightforward. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs, OpenAI warns. At scale, distillation can accelerate the transfer of advanced capabilities without requiring the same investment in safety. The company says these concerns become heightened as models gain capabilities in dual-use domains.

No breach, no broken encryption

The operators did not break OpenAI's encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester — in a coordinated, scaled manner that violated OpenAI's terms of service.

One observed technique stands out for its novelty: operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe the hidden reasoning content. In effect, the attackers turned the model itself into the decryption oracle, sidestepping the encryption that protects reasoning traces in transit.

OpenAI emphasizes that this manipulation is not a vulnerability unique to its models. The company has shared information about it with industry partners through the Frontier Model Forum to strengthen collective defenses against adversarial distillation.

Independent security researchers also played a role in surfacing the attack class. OpenAI says researchers brought related cross-model and conversation-compaction vulnerabilities to its attention through responsible disclosure, and the company investigated their findings and confirmed the attack paths they identified were real. "Their work helped us understand the broader attack class and accelerate mitigations," OpenAI writes.

Attribution: Moonshot AI

OpenAI's attribution statement is deliberately measured. "It is unclear whether all operators we observed during the relevant time period originated from a single actor," the company writes. "However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi."

The phrasing stops short of accusing the company itself. It identifies individuals associated with Moonshot AI as responsible for a core cluster, not the organization's official infrastructure or leadership. The activity also evolved over time, which OpenAI says reinforces that adversarial distillation is a broader security challenge requiring layered, adaptive defenses.

The attribution lands amid an intensifying commercial and geopolitical contest over frontier model capabilities. Reasoning models — which expose or partially expose their chain of thought as part of generating an answer — have become a competitive battleground, and distillation offers a cheaper path to mimicking advanced capabilities than training from scratch. OpenAI frames the risk in national security terms: extracted reasoning transferred to another model can shed the safety guardrails the original developer built.

The response

OpenAI mitigated the campaign through a combination of account enforcement, technical controls, and partner coordination. The company banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks.

On the technical side, OpenAI strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. It closed the pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents — the cross-conversation decryption trick the operators used. It also added checks to detect and hold streamed output that might expose reasoning.

When related activity moved through third-party services, OpenAI worked with those providers to identify and disrupt the accounts involved. The company also shared findings through the Frontier Model Forum and appropriate government information-sharing channels, so that other frontier developers and public-sector partners could look for similar activity and strengthen their own defenses.

OpenAI notes a structural implication: systems that support portable or replayable reasoning artifacts may face related risks. That observation extends beyond OpenAI's own products to any deployment where reasoning traces can be exported, replayed, or fed into another model.

Before publishing, OpenAI says it investigated the scope and potential impact, deployed its own mitigations, and shared with and took feedback from researchers and industry partners to ensure protections against this type of attack are in place. Additional mitigation and investigation work is continuing.

Why this matters beyond OpenAI

The disclosure reframes distillation from a competitive nuisance into a security category with its own threat model. OpenAI explicitly states the risk is not unique to its systems and that similar techniques may affect other advanced AI systems, making this a shared security challenge requiring industry-wide coordination.

The incident also illustrates how reasoning models have changed the attack surface. When models emit hidden or encrypted chain-of-thought as part of their output pipeline, that artifact becomes a target — one that can be copied, replayed, and decrypted using the model itself. Traditional content filtering on visible text does not address this class of attack.

Looking ahead, OpenAI expects adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic their capabilities. The company identifies two gaps in its current defenses: partner-hosted deployments need the same protections as first-party services, and tool-output attacks require protections that examine more than ordinary visible text.

OpenAI's response will continue to focus on three areas: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing across industry and government. The company says it is continuing to improve tool defenses, classifier coverage, model refusals, and to propagate relevant controls across cloud partners.

Original: arxiv.org

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

148 articles

Related articles

  1. OpenAI models broke out of isolation and breached Hugging Face
  2. OpenAI Bans Accounts Reviving Russia's 'Stop News' Influence Operation
  3. OpenAI launches misalignment disclosure framework, publishes six reports
  4. OpenAI Halts Training of Its Most Powerful Models
  5. OpenAI Explains How Its Safety Pipeline Missed GPT-4o Sycophancy

« Previous article