Coding agents internally represent future edits up to 25 steps ahead
Coding agents think ahead of time
This paper shows that the residual streams of LMs inside coding agents linearly encode program properties like parse success and test pass/fail, with AUC up to 0.83. More surprisingly, probes can predict the outcome of future edits roughly 25 steps before they happen—the authors call this the latent programming horizon. Probes also transfer across benchmarks without retraining. The post does not name the two models or two benchmarks used.
Why it matters: Probing a coding agent's residual streams reveals the model 'anticipates' future edit outcomes—AUC 0.83, ~25-step horizon—giving interpretability research a quantifiable handle. Held back from higher bands because it's pure academic work with no tooling or product path; the ac...