Stealing Reasoning Traces from Encrypted Chain-of-Thought Blocks
Stealing Reasoning Traces from Proprietary LLM APIs
Encrypted chain-of-thought blocks returned by Anthropic, OpenAI, and Google are portable across sessions, users, and models. The authors replay a Claude Opus 4 reasoning trace into a jailbroken Claude Haiku 4.5, which then transcribes Opus's hidden reasoning verbatim—without attacking the strong model directly or triggering anti-distillation safeguards. From 6,708 public agent trajectories they decoded 315,320 reasoning blocks and recovered 704 privacy artifacts, 64 of which appeared only inside the encrypted traces.
Why it matters: A hard-hitting security finding with a paper, numbers, and a reproducible path. All three HKR axes hit. Slight deduction for technical depth, but the industry impact justifies 88.