TL;DR
Researchers showed attackers could replay hosted session artifacts, called “reasoning blobs, ” to recover models’ intermediate chains-of-thought on some cloud-hosted endpoints. OpenAI says it detected and shut down a coordinated “adversarial distillation” campaign (banning more than 15, 000 accounts and applying fixes), but the researchers reported the same extraction techniques still worked on certain Microsoft Azure endpoints weeks later (reported by The Decoder and documented by the research team: arXiv:2608.09867 and stolen_thoughts_update.pdf).
Top 3 immediate asks for leaders
- Confirm hosting and crypto controls. Ask vendors where models are hosted and whether intermediate artifacts are bound to per-session ephemeral keys.
- Run adversarial-distillation tests. Include adversarial replay / notepad-style prompts in your AI penetration tests against hosted endpoints within 30 days.
- Contractually require session-binding. Add a procurement clause that requires per-session or per-tenant cryptography and attestation of streamed-output filtering from cloud hosts.
Why this matters
Chains-of-thought are more than verbose logs. They capture a model’s intermediate reasoning, attention patterns, and sometimes training shortcuts or private data. If attackers can recover those internal steps, they can train or prompt cheaper models to behave like pricier ones, what researchers call “adversarial distillation.” That creates a supply-chain risk: protections a model vendor applies at its API can be undermined by differences in how cloud hosts handle sessions, encryption, and streaming.
What happened (concise, technical)
Two related techniques were reported by researchers:
- Encrypted-packet replay + decryption-oracle (reported by researchers). Attackers copied session artifacts, described as encrypted “reasoning blobs” or encrypted session artifacts, from one session and fed them into weaker endpoints or models that acted like a decryption oracle, returning the stronger model’s intermediate reasoning in plaintext. Researchers frame this as letting a weaker model learn from a stronger one.
- Tool-notepad trick (public demo by Can Bölük, documented by researchers). A prompt and tool pattern coerces a model to write intermediate steps into an accessible tool buffer or “notepad.” If that tool buffer is readable by the user, it exposes the intermediate reasoning. Researchers reported this worked against several models in their tests and failed against others (see their update for model-by-model details).
“We stole reasoning. Again.”, Joachim Schaeffer (researcher), X post
Important technical nuance: researchers describe the recovery as a model behavior that prints or reconstructs the intermediate steps when given those artifacts, not a classical cryptographic break. The root causes they highlight include session artifacts being replayable or session keys not being tightly bound to a single session or tenant.
Compact timeline (reported)
- Activity began at low volume on July 1, 2026 (reported by OpenAI).
- Spike on July 24-25 to roughly 16, 000 requests from more than 4, 000 users; OpenAI identified a network of over 15, 000 related accounts and said it “fully shut down” those accounts by July 28 (OpenAI reporting summarized in The Decoder, Oct 1, 2026).
- Researchers tested hosted endpoints on September 13 and reported the techniques were blocked on OpenAI’s and Anthropic’s own APIs but still worked against models served on Microsoft Azure, including GPT‑6 Astra and Anthropic Sonnet 5 (reported by the researchers; see arXiv:2608.09867 and stolen_thoughts_update.pdf).
- Researchers reported that OpenAI added safeguards to the Azure endpoint on September 27 and that Anthropic’s fixes were effective on September 28 (per the researchers’ update).
Who said what (attribution)
- OpenAI characterized the activity as an “adversarial distillation” campaign and reported account bans and platform mitigations (OpenAI statements summarized in reporting by The Decoder).
- Joachim Schaeffer and coauthors documented the attack class and experiments in a paper available as arXiv:2608.09867 and in a follow-up project update (stolen_thoughts_update.pdf); Schaeffer posted summaries on X.
- The Decoder (Maximilian Schreiner, Oct 1, 2026) compiled vendor statements and the researchers’ findings into public reporting.
- OpenAI said it tied a core group of accounts to people associated with Moonshot AI (maker of the Kimi model) but noted attribution was not conclusive. Treat that as vendor attribution unless Moonshot AI or independent auditors confirm otherwise.
What enterprise leaders should ask vendors and cloud hosts
Don’t accept “we host it for you” as sufficient. Ask these specific technical and contractual questions:
- Are intermediate model artifacts (streamed tokens, tool outputs, session buffers) bound to per-session ephemeral keys?
- Do you use per-tenant or hardware-backed keys (TPM/TEE) to isolate encryption domains across customers?
- Are streamed outputs filtered server-side for chain-of-thought patterns? Can you provide audit logs or attestations of that filtering?
- Do you permit third-party or customer-run tools to access intermediate buffers? If so, how are those tool channels sandboxed?
- Can you supply an incident timeline and evidence when a hosting vulnerability is reported that affects model confidentiality or internal reasoning?
Sample SLA clause starters you can adapt:
- “Provider shall bind all intermediate model artifacts to per-session ephemeral keys and provide attestation of session-binding on request.”
- “Provider shall provide quarterly audit evidence demonstrating streamed-output filtering for chain-of-thought patterns and allow vendor-approved adversarial tests.”
Mitigations that help and their trade-offs
- Per-session ephemeral keys. Strength: prevents blobs being replayed across sessions. Trade-off: added key-management complexity during autoscaling and multi-region deployments.
- Per-tenant / hardware-backed keys. Strength: isolates tenants cryptographically. Trade-off: requires hardware support and can add latency and cost.
- Streamed-output screening with heuristics. Strength: can block obvious chain-of-thought leakage in real time. Trade-off: pattern-based filters produce false positives and negatives and can undermine explainability features.
- Tool sandboxing and write-only buffers. Strength: prevents tools from exposing intermediate buffers to users. Trade-off: may break legitimate plugin or developer workflows that rely on persisted intermediate state.
- Regular adversarial testing and attestation. Strength: validates defenses against realistic attacks. Trade-off: requires investment in red-team capability and legal guardrails for testing third-party-hosted models.
No single control is foolproof; layered defenses plus regular testing are essential.
Key takeaways, quick Q&A
-
Did OpenAI stop the campaign?
OpenAI reported detecting the “adversarial distillation” activity, banning more than 15, 000 related accounts and applying platform fixes (as summarized in reporting by The Decoder). Those actions addressed the activity observed on OpenAI’s own API surface.
-
Could attackers actually extract internal reasoning?
Researchers (Joachim Schaeffer et al., arXiv:2608.09867 and the stolen_thoughts_update.pdf) demonstrated methods that in their tests recovered chains-of-thought verbatim from certain hosted endpoints; they report some single-attempt successes in controlled experiments.
-
Were cloud endpoints still vulnerable after vendor patches?
According to the researchers’ September tests, Azure-hosted endpoints were vulnerable on Sept 13 even after vendors patched their own APIs; the researchers reported Azure mitigations were added on Sept 27. These timeline claims come from the research team’s papers and updates and should be confirmed with Azure/OpenAI public statements for enterprise risk assessments.
-
Does this prove IP theft at scale?
The publicly reported shutdowns reflect attempted extractions. The materials note these were attempts and do not conclusively demonstrate large-scale, durable distillation or reproduction of vendor models in production. Organizations should treat the incidents as a serious proof-of-concept risk rather than confirmed widescale IP theft without further evidence.
-
What immediate steps should buyers of hosted models take?
Confirm hosting locations and cryptographic controls, require session-binding and per-tenant crypto in contracts, run adversarial-distillation tests on critical endpoints, and demand incident timelines and attestations from cloud hosts.
Where the story still needs clarity
- Exact root causes in each vulnerable hosting setup: whether the issue was shared encryption keys, session-binding flaws, or streaming/tool integration problems (the research papers and vendor statements are the primary artifacts for this).
- Whether attackers achieved durable, production-scale distillations that fully reproduced targeted models’ capabilities, reported activity includes attempts, but large-scale success is not confirmed in public materials.
- Vendor and cloud-provider confirmation of dates and mitigations. The researchers report Azure protections added on September 27; enterprises should seek Microsoft/Azure and OpenAI statements for their procurement risk files.
- Attribution to named entities: OpenAI tied some activity to accounts linked to people associated with Moonshot AI but said attribution was not conclusive. Treat that as vendor attribution pending independent verification.
Final thought for leaders
Boards and CISOs should add hosting-chain security to their AI risk registers, insist on auditable session-binding and attestation from vendors and hosts, and treat adversarial-distillation testing as a standard part of AI security validation.