Skip to content
Synced · WeChat

ACL 2026 Survey: Intrinsic Interpretability Moves LLMs from Post-hoc Analysis to Design

ACL 2026 综述:从事后解释到内生解释,大模型内生可解释性的前沿进展

ACL 2026 Main accepted a survey on intrinsic interpretability for LLMs, grouping methods into five design paradigms. It covers functional transparency, concept alignment, decomposable representations, explicit modularization, and latent sparsity induction, with MoE, CBM, and GLU/SwiGLU examples. The key test is whether interpretable parts sit on the model’s computation path, not outside it.

Why it matters: HKR-H/K/R pass: the survey has a clear framing shift, five named mechanisms, and safety/debugging relevance. It is a useful research release, not a model launch or empirical breakthrough.

Read the original ↗Export Markdown