When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

Tobias Bersia, Tatiana Gaintseva

arXiv:2607.23379 · 2026-08-10 공개 · arXiv · PDF

fine-tuning model-interpretability activation-oracles hidden-concepts layer-ablation behavioral-leakage readout-pathway taboo-word-guessing

Abstract

Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.

한국어 요약

한 줄 요약

AO가 훈련된 개념을 의도적으로 회피하는 현상을 밝혀내며, 학습된 해석 도구의 신뢰성 문제를 제기한다.

핵심 기여도

핵심 아이디어

Activation Oracles(AO)는 다른 모델의 내부 활성화 정보를 자연어 질문에 대답하는 방식으로 해석하는 언어 모델이다. 이 연구는 AO가 학습된 시스템이기 때문에, 단순히 정보를 읽어내는 도구가 아닌, 훈련 데이터와 목적에 따라 정보를 선택적으로 표현하거나 회피할 수 있음을 제시한다. 연구는 Taboo Word Guessing 설정에서 주제 모델이 내부적으로 특정 단어(예: leaf, moon)를 사용하면서 외부 표현은 회피하도록 fine-tuning 되는 상황을 조사한다. AO가 이 주제 모델의 활성화 정보를 기반으로 훈련되었음에도 불구하고, 해당 단어를 회피하는 현상을 관찰했다. 이는 AO가 단순히 정보를 읽는 것이 아니라, 훈련 과정에서 개념별 보고 정책을 학습할 수 있음을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용

이 연구는 AI 모델의 내부 정보를 해석하는 AO 기반 도구의 신뢰성 문제를 제기하며, 모델 감사, 보안, 투명성 검증 등에서 AO의 한계를 인지하고 보완 전략을 수립하는 데 활용될 수 있다. 특히, 보안 모델이나 비공개 정보를 다루는 시스템에서 AO의 선택적 보고 정책이 정보 누출을 방지하는 도구로 활용될 수 있다.