Skip to content
Hacker News front page

LLM language outputs are unreliable for security monitoring

The Implications of Linguistic Illegibility for LLM Security

James Mickens introduces 'linguistic illegibility': an LLM's text outputs or mechanistically extracted language features may not reflect its actual internal computation. Since models do math over activation spaces and only translate to language at the input/output ends, the translation is lossy. This means security mechanisms that rely on linguistic self-reporting—chain-of-thought monitoring, constitutional self-critique, activation probing with language-defined features—can never be fully sound. He argues for sandbox isolation that doesn't depend on reading linguistic state at all, proposing taint tracking to define system state that must never be influenced by model outputs, plus robust virtualization and third-party auditing. The post does not include experimental data; it's a conceptual argument and design proposal.

Why it matters: Mickens unifies CoT monitoring, self-reflection, and activation probing under one theoretical vulnerability: internal computation happens in activation space, with lossy translation to language only at the endpoints. Strong concept, but a pure theory paper with no empirical va...

Read the original ↗Export Markdown