Safety

AI-generated text

OpenAI disrupted coordinated campaign to extract models' hidden reasoning

OpenAI says it identified and disrupted a coordinated campaign, beginning in early July, that attempted to extract protected internal reasoning from its models — a technique known as adversarial distillation.

OpenAI disrupted coordinated campaign to extract models' hidden reasoning

OpenAI detected and disrupted a coordinated campaign aimed at extracting protected internal reasoning from its models. The company says the earliest observed activity dates to the first week of July, and that it investigated the scope and potential impact before publishing its findings.

Definitions: protected reasoning and adversarial distillation

  • "Protected reasoning" refers to a model’s internal chain-of-thought or working record used to solve a task — information that may be withheld from the model’s final user-facing answer.
  • "Adversarial distillation" describes the systematic, unauthorized use of one model’s outputs or internal reasoning to train, reproduce, or improve another model. Extracting protected reasoning can enable others to reproduce a model’s capabilities without preserving the original safety constraints.

How the campaign worked — not a system breach

OpenAI states the operators did not break encryption, compromise a database, or gain direct access to stored user conversations. Instead, attackers manipulated model interactions so protected reasoning could be reproduced in forms visible to requesters. This was done in a coordinated, scaled way that violated OpenAI’s terms of service. The company notes this manipulation is not a vulnerability unique to its models and has shared information with industry partners via the Frontier Model Forum to strengthen collective defenses.

Timeline and scale

  • Activity began on July 1, initially at low volume.
  • High-volume spikes occurred on July 24–25: roughly 16,000 requests using relevant extraction patterns came from over 4,000 users on those days.
  • Further investigation revealed related prompt-pattern activity across a cluster involving more than 15,000 users; OpenAI says it fully disrupted this activity by July 28.

Role of independent security researchers

Independent security researchers responsibly disclosed related cross-model and conversation-compaction vulnerabilities to OpenAI. The company investigated and confirmed the attack paths identified by those researchers, and their work helped accelerate mitigations and improve understanding of this broader attack class.

Attribution

OpenAI says it is unclear whether all observed operators belonged to a single actor, but attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.

Why this matters: risks

Extracted reasoning can be used to train other models without the safety safeguards applied to the original model’s user-facing outputs. At scale, adversarial distillation can accelerate the transfer of advanced capabilities without equivalent investment in safety, increasing concerns in dual‑use domains. OpenAI emphasizes this is an industry‑wide security challenge, not one limited to its systems.

Mitigations deployed

OpenAI describes a mix of actions it took to mitigate the campaign:

  • Account enforcement: banning or restricting fraudulent accounts.
  • Strengthened signup and infrastructure controls.
  • Expanded monitoring for related networks.
  • Reinforced protections for hidden reasoning across users, workspaces, organizations, and model families.
  • Closed an attack pathway that allowed replay of another user’s encrypted reasoning to recover its contents.
  • Added checks to detect and hold streamed outputs that might expose reasoning.
  • Coordinated with third‑party service providers to identify and disrupt accounts when related activity moved through those services.

Information sharing and coordination

OpenAI shared relevant findings through the Frontier Model Forum and appropriate government information‑sharing channels so other frontier developers and public‑sector partners could look for similar activity and strengthen defenses. The company warns that systems supporting portable or replayable reasoning artifacts may face related risks.

Ongoing work and outlook

OpenAI expects adversarial distillation attempts to become more sophisticated as frontier models improve and actors seek cheaper ways to mimic capabilities. Defending against this activity requires layered controls and continual adaptation. The company is continuing work to improve tool defenses, classifier coverage, model refusals, and propagation of relevant controls across cloud partners.

Final note

OpenAI stresses that the reported figures describe attempted extractions, not necessarily successful ones, and that mitigation and investigation efforts are ongoing.