OpenAI Shuts Down 15,000-Account Reasoning Theft Campaign, but Azure Ran It 61 More Days
Researchers confirm the same encrypted-reasoning flaw kept working on Azure after OpenAI patched its own API

OpenAI disclosed that it detected and disrupted a coordinated campaign in which more than 15,000 fraudulent accounts — tied to individuals associated with Chinese AI company Moonshot AI — systematically extracted the hidden reasoning of its frontier models through a flaw in how encrypted reasoning tokens are shared across sessions. The attack peaked at 16,000 extraction requests on July 24 and 25 before OpenAI neutralized it on July 28. OpenAI notes these figures describe attempted, not necessarily successful, extractions. What the company did not publicize until now: the identical attack continued working on Microsoft Azure's hosted version of the same OpenAI models for another 61 days, through September 27, according to independent security researchers who retested the vulnerability after OpenAI's own patch.
The disclosure coincides with a September update from Joachim Schaeffer's research team at the MATS Program and ELLIS Institute Tübingen, which first documented the underlying vulnerability in a paper submitted to arXiv on August 10. The combination means an enterprise customer routing traffic through Azure OpenAI Service during those nine weeks could have believed they were protected — based on OpenAI's own timeline — while their actual exposure continued through an intermediary endpoint Microsoft controls independently.
The Flaw in Encrypted Reasoning Blocks
When a reasoning-capable model like GPT-6 Astra or Claude Opus returns a response via API, the provider also returns an encrypted opaque blob containing the model's hidden chain-of-thought. This blob must be echoed back in each API call for stateless conversation management — providers marketed it as protection for proprietary reasoning. The authenticated encryption provides confidentiality and tamper resistance but critically lacks binding to a specific session, user, or model. An encrypted reasoning block produced by GPT-6 Astra is accepted as valid input by GPT-5.6 Luna; a Claude Opus 4.8 block is processed by Haiku 4.5. The encryption was not cracked — it was bypassed by routing the block through a cheaper sibling model acting as an unwitting decoder, as Panfilov et al. document in detail.
The attack chain is three steps: query a capable target model to capture its encrypted reasoning block; submit that block to a weaker model from the same provider; instruct the weaker model to "transcribe prior thoughts verbatim." The weaker model complies — its refusal training guards against disclosing its own reasoning, not replayed content from a different model in the provider family. Confirmed decoder models include Haiku 4.5 for Claude traces, GPT-5.6 Luna for OpenAI traces, and Gemini Robotics 1.6 for Google. The research team estimated a cost of roughly $720 to decrypt 10,000 reasoning traces — economically trivial for a funded AI laboratory.
The July campaign targeted this architectural gap at scale specifically on reasoning models. GPT-6 Astra, released on September 3, is built around what OpenAI describes as a "recurrent depth" or "looped transformers" approach that enhances reasoning efficiency but, by design, obscures some or all of the model's chain-of-thought from external inspection. Ironically, the more a provider obscures reasoning from the visible output, the more valuable the encrypted trace becomes to an adversary who can extract it — the hidden process is precisely what those higher-tier models are trained on and what a competing lab most wants.
A Second Attack Required No Encryption at All
Schaeffer's team also confirmed a simpler technique: instruct a model to write its reasoning into a "virtual notepad" tool, a standard capability in agent frameworks, and read whatever the model wrote. This approach required no encrypted-block manipulation and worked against every tested OpenAI and Anthropic model with three exceptions: Opus 5, Fable 5, and Fable 5.1. Those three Anthropic models appear to carry additional tool-use refusal training. Fable 5 and Fable 5.1 are Anthropic's safety-hardened tiers — Fable 5.1 with enhanced controls for biology, cybersecurity, and LLM research — and their immunity likely reflects deliberate hardening for high-risk domains rather than architectural accident.
Separately, the Panfilov team's scan of public GitHub repositories and Hugging Face agent datasets found 315,320 encrypted reasoning blocks in shared logs that developers had inadvertently included. After applying the decryption technique, they recovered 182 credentials — including 62 API keys, 33 passwords, and 24 access tokens — and 367 PII artifacts, with 64 items appearing only inside the encrypted reasoning, never in visible plaintext. Developers who strip sensitive data from conversation logs before publishing may still be leaking the same data inside the reasoning blob they included alongside it.
Read more: OpenAI's persistent Dots agents and the pattern of safety incidents behind them
The Azure Gap: Same Models, Different Protections
OpenAI patched its own API on July 28. Azure OpenAI Service — Microsoft's resale of OpenAI's model API — ran on Microsoft's separate patching timeline. On September 13, Schaeffer's team retested against Azure and found every tested OpenAI model, including GPT-6 Astra (released just ten days earlier on September 3), still vulnerable. Anthropic models accessed through Azure were also extractable through Claude Sonnet 5. The fix landed on Azure on September 27; Anthropic's models on Azure were protected from September 28, according to the researchers' account as reported by WindowsForum.
"Same models, but different protections depending on which platform serves them," Schaeffer summarized. Microsoft has not publicly responded.
This gap is not a coordination anomaly — it is a structural feature of the reseller model. Azure OpenAI Service controls its own endpoint infrastructure; a provider's security fix does not automatically propagate to the reseller. Enterprise customers who verify security posture against the primary provider — a reasonable assumption — may be exposed through an intermediary on an independent patching schedule. A 61-day gap is enough time for a motivated, well-funded actor to complete an extraction campaign that the enterprise believed was already blocked.
Read more: Microsoft's AI rulebook bars models from hiding reasoning — while Azure lagged on this patch
Moonshot AI and the Broader Attribution
OpenAI attributes the core cluster of the July campaign to individuals associated with Moonshot AI, the Beijing-based developer of the Kimi model series. Founded in March 2023 by Tsinghua University alumni Yang Zhilin, Zhou Xinyu, and Wu Yuxin, Moonshot reached a $35 billion valuation in July 2026 and is backed by Alibaba and Tencent. OpenAI notes uncertainty: "it is unclear whether all operators we observed during the relevant time period originated from a single actor."
Moonshot has faced prior accusations of distillation. In September 2026, Anthropic accused the company of routing user requests through Claude models rather than processing them using Kimi, and retaining a subset of those exchanges for training purposes (incident GTG-16002). In April 2026, a US congressional subpoena targeted companies using Kimi models. Moonshot has not responded to media requests regarding the OpenAI disclosure.
The timing of July's extraction campaign is notable. Moonshot released Kimi K3 — a 2.8 trillion-parameter open-weights model the company benchmarked as competitive with US frontier models — on July 16, 2026, with its full weights published July 27. The attack peaked on July 24 and 25, about eight days after the Kimi K3 launch and two to three days before the full weight release. No direct connection between any extracted reasoning and Kimi K3's training has been publicly established.
Why Reasoning-Trace Extraction Differs from Standard Distillation
Standard adversarial distillation generates large volumes of question-and-answer pairs from a competitor's API and uses those outputs to train a student model. That captures response style and factual patterns but not the underlying reasoning process. Reasoning-trace extraction targets the hidden chain-of-thought directly — the step-by-step deliberation that is the product of expensive reinforcement learning and post-training procedures. A September 2026 CISA advisory on Chinese AI distillation campaigns noted that current export control frameworks restrict trained model weights, yet a determined actor can approximate a substantial share of a frontier model's value without ever obtaining the underlying weights. Direct reasoning-trace extraction goes further still: it requires no weight transfer and no large-scale API mining, just a handful of sessions to capture the reasoning process itself.
Policy analysts have called for Entity List designations for Moonshot AI, DeepSeek, and MiniMax, and for executive action targeting proxy operators. Thursday's OpenAI disclosure provides the most direct public attribution yet of a reasoning-trace extraction campaign to a named company — and the Azure gap adds a dimension that export-control frameworks were not designed to address: the same attack can continue working through a licensed American cloud provider after the primary model vendor has already applied its fix.
For organizations that used Azure OpenAI or Azure-hosted Anthropic models between July 28 and September 27, it is worth reviewing whether sensitive data was processed through agentic reasoning steps during that period. Both endpoints are now patched. Developers building agentic pipelines should treat encrypted reasoning blobs the same way they treat cleartext secrets — stripped from shared logs before publication, not echoed into open repositories alongside the rest of the API transcript.
OpenAI said it shared its findings through the Frontier Model Forum and applied "account enforcement, technical controls, and partner coordination" — banning fraudulent accounts and expanding monitoring for similar patterns. Whether that coordination produces synchronized patch timelines for cloud resellers on future vulnerabilities is the question the Azure gap raises most directly. The current reseller model means any enterprise customer who purchases AI model access through a cloud intermediary is implicitly trusting that intermediary's patching discipline as much as the primary provider's — a dependency that was invisible until a 61-day discrepancy made it concrete. As frontier models are distributed through more cloud platforms, managed API services, and regional resellers, the surface area for that gap grows.