科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACM Computing Surveys2026-01-23· Interpretability

Bridging the Black Box: A Survey on Mechanistic Interpretability in AI

Shriyank Somvanshi, Md Majharul Islam, Amir Rafe, Anannya Ghosh Tusti, Arka Chakraborty, Anika Baitullah, T Chowdhury, Nawaf Alnawmasi, Anandi Dutta, Subasish Das

原始摘要(英文原文)· Original abstract
Mechanistic interpretability seeks to reverse-engineer the internal logic of neural networks by uncovering human-understandable circuits, algorithms, and causal structures that drive model behavior. Unlike post hoc explanations that describe what models do, this paradigm focuses on why and how they compute, tracing information flow through neurons, attention heads, and activation pathways. This survey provides a high-level synthesis of the field-highlighting its motivation, conceptual foundations, and methodological taxonomy rather than enumerating individual techniques. We organize mechanistic interpretability across three abstraction layers— neurons , circuits , and algorithms —and three evaluation perspectives: behavioral , counterfactual , and causal . We further discuss representative approaches and toolchains that enable structural analysis of modern AI systems, outlining how mechanistic interpretability bridges theoretical insights with practical transparency. Despite rapid progress, challenges persist in scaling these analyses to frontier models, resolving polysemantic representations, and establishing standardized causal benchmarks. By connecting historical evolution, current methodologies, and emerging research directions, this survey aims to provide an integrative framework for understanding how mechanistic interpretability can support transparency, reliability, and governance in large-scale AI.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Bridging the Black Box: A Survey on Mechanistic Interpretability in AI — 科研速览 Science Skim