Agent Security Has No Universal Sandbox: A Five-Layer Interception Chain from Intent to Outcome
Using a case where npm test hides a malicious subprocess that steals SSH keys, this piece breaks Agent security into five layers: internal activation probing, chain-of-thought auditing, structured tool-call authorization, static command inspection, and kernel-level sandboxing plus resource-side immutable boundaries. The core insight: layers closer to the model understand intent better but are easier to bypass; layers closer to the OS enforce hard limits but understand zero business semantics. The post explicitly states that J-lens mind-reading is only probabilistic early warning, chain-of-thought can be unfaithful, the MCP gateway can't see dynamic subprocesses spawned by npm test, and static rules miss runtime child processes—only Landlock/seccomp or gVisor/Firecracker isolation finally blocked the exfiltration. It also debunks three sandbox myths: cutting the network doesn't make you safe, VMs can leak cloud credentials via metadata services, and detect-and-kill loses the race to exfiltration.
Why it matters: The five-layer interception framework is original, with concrete techniques and failure modes at each layer — not generic security fluff. Held back from 85 because the article body is truncated mid-argument, missing the full reasoning and deployment examples.