When thousands of accounts ask variations of the same unusual question, an abuse problem can become an extraction operation. OpenAI says that is what it found in a coordinated adversarial-distillation campaign disclosed on September 30: operators tried to make its systems reproduce reasoning that would normally remain hidden.

The company says the activity began on July 1 and surged on July 24 and 25, when it recorded 16,000 requests using a relevant extraction pattern across more than 4,000 users. A broader investigation connected related prompt patterns to a cluster exceeding 15,000 users. OpenAI says it fully disrupted that cluster by July 28.

Two qualifications matter. Those figures count attempted extractions, not necessarily successful ones. OpenAI also says the operators did not break encryption, compromise a database, or directly access stored conversations. Instead, they manipulated interactions between models in an effort to turn protected reasoning artifacts into readable output.

Distillation is more than copying answers

Model distillation can be legitimate when one system learns from another model's outputs with authorization. It becomes adversarial when collection is designed to reproduce capabilities or intellectual property without consent and evade a provider's controls.

In the campaign OpenAI describes, operators copied encrypted reasoning from one conversation and asked a model in another conversation to decode and transcribe it. OpenAI attributes a core cluster to people associated with Moonshot AI, the developer of Kimi, while acknowledging that it cannot determine whether every observed operator belonged to one actor. That attribution is the company's assessment, not a publicly proven independent finding.

Scale changes the defense model: request-level filters still matter, but they are insufficient when an operation spreads across accounts, models, and third-party services.

Independent researchers had already documented a related class of weaknesses in the paper “Stealing Reasoning Traces from Proprietary LLM APIs”. They found that encrypted blocks reusable across sessions or models could enable reasoning extraction and other attacks. OpenAI says it confirmed the paths disclosed by the researchers and used their responsible report to accelerate mitigations.

Portable reasoning creates a larger boundary

OpenAI's response combined account restrictions, stronger signup controls, network monitoring, and checks designed to hold streamed output that might expose reasoning. It also closed a replay path for encrypted artifacts and shared relevant indicators through the Frontier Model Forum and government channels.

For teams building agents, the practical lesson extends beyond one provider. Reasoning artifacts that move among sessions, organizations, model families, or hosted services create a broader attack surface. An opaque block is not inherently safe if another component can interpret it outside the original context and permission boundary.

Public evidence is not sufficient to quantify how much useful knowledge, if any, was transferred through this campaign. The operational account comes from OpenAI, and the company has not published all evidence needed for independent verification. Even with that limitation, the incident points to a clear engineering shift: protecting advanced models requires aggregate behavior detection, tighter controls over portable artifacts, and equivalent safeguards across first-party and partner deployments. Conversation-by-conversation security is too small for a coordinated campaign.

Sources: OpenAI's technical account and the independent researchers' paper.