Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Samuel Marks, C. Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller

arXiv:2403.19647 · 2026-07-27 공개 · arXiv · PDF

language-models interpretable-ai model-behavior causal-graphs feature-ablation sparse-feature-circuits shift-method unsupervised-interpretability

Abstract

We introduce methods for discovering and applying sparse feature circuits. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors. Circuits identified in prior work consist of polysemantic and difficult-to-interpret units like attention heads or neurons, rendering them unsuitable for many downstream applications. In contrast, sparse feature circuits enable detailed understanding of unanticipated mechanisms. Because they are based on fine-grained units, sparse feature circuits are useful for downstream tasks: We introduce SHIFT, where we improve the generalization of a classifier by ablating features that a human judges to be task-irrelevant. Finally, we demonstrate an entirely unsupervised and scalable interpretability pipeline by discovering thousands of sparse feature circuits for automatically discovered model behaviors.

한국어 요약

한 줄 요약

"스파스 피처 서킷"을 통해 언어 모델 내 인터프리터블한 인과 그래프를 자동으로 발견하고 편집하는 방법을 제시한다.

핵심 기여도

핵심 아이디어

기존 연구는 어텐션 헤드나 MLP 모듈과 같은 다의성 유닛을 사용해 모델 행동을 설명했으나, 이는 해석이 어려워 실용적 활용이 제한적이었다. 본 연구는 **인간이 해석 가능한 단일 의미의 피처**(interpretable features)를 기반으로, **인과성을 가진 서킷**(causal circuits)을 발견하는 새로운 접근법을 제시한다. 이를 위해 **스파스 오토인코더**(SAE)를 사용해 잠재 공간 내 인터프리터블한 방향을 학습하고, **선형 근사**(linear approximation)를 통해 인과성을 효율적으로 추정한다. 이는 기존의 인과 추정 방법(예: attribution, integrated gradients)보다 계산 효율성이 높으며, **IE (intervention effect)**를 기반으로 노드와 엣지의 중요도를 추정한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 언어 모델 내 인터프리터블한 인과 구조를 **자동으로**, **비지도 학습 기반으로** 발견하는 데 기여한다. 특히, **인간 판단 없이도 수천 개의 행동에 대한 서킷을 생성**할 수 있다는 점에서 실용적 가치가 크다. 또한, **SHIFT 기법**은 모델의 불필요한 신호에 대한 민감도를 제거함으로써 일반화 성능을 향상시키는 데 효과적이다.

한편, **IE 추정의 정확도는 early layer에서 낮을 수 있으며**, 이는 선형 근사의 한계로 인한 것이다. 또한, **모든 모델 행동에 대한 서킷이 자동으로 발견되지는 않으며**, 일부 행동은 해석이 어려울 수 있다.

실용적 활용