AI Affairs, home

Sunday 4 October 2026

Technology

OpenAI says it disrupted a campaign to extract protected model reasoning

The company linked a core cluster of activity to individuals associated with Moonshot AI and described how encrypted reasoning could be exposed across conversations.

The Pioneer Building in San Francisco, partly screened by trees along a street.
Photo: HaeB, CC BY-SA 4.0, via Wikimedia Commons (cropped)

OpenAI said on 30 September 2026 that it had disrupted a coordinated campaign to extract protected reasoning from its models. OpenAI linked a core cluster of the activity to people associated with Moonshot AI, which develops Kimi, and said operators shaped model exchanges to expose hidden reasoning to requesters.

Key points

  • OpenAI observed 16,000 requests using a relevant extraction pattern from over 4,000 users; it says the figures count attempts, not necessarily successful extractions.
  • The company said one route involved moving encrypted reasoning between conversations and asking a model to render its contents as text.
  • OpenAI said it disrupted a related cluster of more than 15,000 users by 28 July and closed a route for replaying another user’s encrypted reasoning.

Encrypted reasoning crossed conversations

Protected reasoning is the internal record a model uses to work through a task, OpenAI said. It can contain material withheld from the final answer. The company describes adversarial distillation as systematically using a model’s outputs or reasoning without authorisation to help train, reproduce or improve another model, and says the campaign it detected was consistent with that practice.

In one technique OpenAI observed, operators copied encrypted reasoning from a conversation into another conversation, then asked a model to decrypt it and write out the hidden content. Think of the encrypted record as a sealed note: possessing it is different from reading it. OpenAI said it closed a pathway through which someone who already had another user’s encrypted reasoning could replay that record and recover what was inside.

Writing an email could produce a final answer that leaves out some of the reasoning used to prepare it. If the protected reasoning from that exchange were recoverable through another conversation, material withheld from the answer could appear as readable text there instead.

OpenAI said the operators did not break its encryption, compromise a database or obtain direct access to stored user conversations. Its account instead describes an attack on the way model interactions handled protected material once an operator could bring an encrypted record into a request. The company said the manipulation violated its terms of service and was carried out in a coordinated manner at scale.

July 24 and 25 brought 16,000 requests

OpenAI said the activity began on 1 July at low volume. On those two days, OpenAI said it recorded spikes of 16,000 requests using a relevant extraction pattern from over 4,000 users. Those are counts of attempted extractions, the company cautioned, rather than a count of reasoning records successfully recovered.

Further investigation found related prompt-pattern activity across a cluster of more than 15,000 users, which OpenAI said it had fully disrupted by 28 July. The broader cluster is a count of users associated with related patterns, while the 16,000 figure counts requests during the two days of high-volume activity. OpenAI said the activity changed over time.

Independent security researchers also reported related vulnerabilities involving interactions across models and the compaction of conversations. OpenAI said it investigated those findings and confirmed that the attack routes the researchers identified were real. The company said their work helped it understand the wider class of attacks and speed up its mitigations.

Moonshot AI attribution covers a core cluster

OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, which develops Kimi. The company said it was unclear whether every operator it observed during the relevant period came from a single actor. Its attribution is therefore narrower than an attribution of the entire cluster of more than 15,000 users to Moonshot AI.

OpenAI said extracted reasoning could help train another model without carrying over safeguards applied to the original model’s visible answers. It also warned that distillation at scale could transfer advanced capabilities without the same safety investment. Those are risks the company identifies for this kind of extraction, rather than measured outcomes of the requests it counted in July.

OpenAI closed a replay route

OpenAI said its response combined account restrictions, technical changes and coordination with partners. It banned or restricted fraudulent accounts, tightened signup and infrastructure controls, and expanded monitoring for related networks. Where activity passed through third-party services, the company said it worked with those providers to identify and disrupt the accounts involved.

The company also said it strengthened protections for hidden reasoning across users, workspaces, organisations and model families. Alongside closing the replay pathway, it added checks intended to detect and hold streamed output that might expose reasoning. OpenAI said partner-hosted deployments need the same protections as its own services, while attacks carried through tool output require checks beyond ordinary visible text.

OpenAI shared its findings through the Frontier Model Forum and government information-sharing channels so that other developers and public-sector partners could look for similar activity. It said systems that allow reasoning records to be moved or replayed may face related risks, and that it is continuing work on tool defences, detection systems, model refusals and controls across cloud partners.

Topics: Foundation models, Safety