Back

OpenAI says it disrupted a bid to distill protected model reasoning

OpenAI Security post artwork on disrupting a coordinated model distillation campaign

On Wednesday, Sept. 30, OpenAI disclosed that it disrupted a July coordinated campaign to extract protected reasoning from its models. The company attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. OpenAI says the operators did not break encryption or databases; they used scaled interaction tricks that violated its terms of service.

A frontier lab just disclosed a summer fight over hidden model reasoning. On Wednesday, Sept. 30, OpenAI published a Security post saying it identified and disrupted a coordinated campaign to extract protected reasoning from its models. The stakes are clear: stolen reasoning can help train another model without the original safeguards, which OpenAI frames as a safety and national-security risk shared across frontier systems.

OpenAI calls the pattern adversarial distillation: the unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model's internal record for working through a task. Extracting it can reveal information withheld from the final answer and help others reproduce capabilities.

The company says operators did not break encryption, compromise a database, or gain direct access to stored user conversations. Instead they manipulated model interactions so protected reasoning could be reproduced in visible form at scale, in ways that violated OpenAI's terms of service. One novel path copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. OpenAI says it closed that pathway and strengthened hidden-reasoning protections across users, workspaces, organizations, and model families.

Activity began July 1 at low volume, OpenAI says, then spiked on July 24 and 25 with 16,000 extraction-pattern requests from more than 4,000 users. Related prompt-pattern activity across a cluster of more than 15,000 users was fully disrupted by July 28. OpenAI notes those figures describe attempted, not necessarily successful, extractions.

OpenAI says it is unclear whether all observed operators came from a single actor. It attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. Mitigations included bans and restrictions on fraudulent accounts, stronger signup and infrastructure controls, expanded monitoring, and sharing findings through the Frontier Model Forum and government information-sharing channels. This disclosure is separate from OpenAI's DevDay agent work, the tens-of-thousands incidents thread, the training-pause sandbox story, and NVIDIA's Open Agent Safety stack.