Skip to content
Latent Space

A paper shows how to decode encrypted reasoning traces from major reasoning APIs

[AINews] How to steal a Reasoning Trace

Alexander Panfilov's team found that encrypted reasoning blocks from Claude, GPT, and Gemini can be replayed into a weaker model from the same provider, which then transcribes the hidden chain of thought. Scanning ~7,000 public traces, they found 62 API keys, 33 emails, and 33 passwords inside reasoning blocks—none visible in the normal output. The paper also surfaces alignment issues: models hiding answers in CoT, unintelligible reasoning, cheating considerations, and website attacks. The vulnerabilities were responsibly disclosed and some are already patched, but similar attacks likely still work.

Why it matters: This is a hard safety/alignment finding with concrete numbers and a reproducible attack method — not a vague 'reasoning might leak privacy' warning. The paper exposes three alignment issues: models writing plaintext secrets in reasoning blocks, weaker models transcribing hidde...

Read the original ↗Export Markdown