Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Markus J. Buehler

arXiv:2607.20058 · 2026-07-26 공개 · arXiv · PDF

language-models hidden-states materials-science open-weight-models mechanism-identification counterfactual-benchmarks state-transformations jacobian-lenses

Abstract

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

한국어 요약

한 줄 요약

Gemma-4-E4B-it 모델에서 물리적 메커니즘 정보가 세 가지 형태로 분리 가능하며, Jacobian lens와 직접 읽기 기법을 통해 내부 표현을 분석하고 제어할 수 있음.

핵심 기여도

핵심 아이디어

기존의 LLM은 정확한 답변을 제공하더라도 내부 표현이 물리적 메커니즘을 반영하는지 여부를 밝히지 못한다. 본 연구는 Gemma-4-E4B-it 모델에서 물리적 메커니즘 정보가 세 가지 실험적으로 분리 가능한 형태로 존재함을 보인다. 첫째, 개별 hidden state에서 개념이 읽을 수 있으며, 둘째, 상태 간의 제어된 변환으로는 구성적 방향성이 전달되고, 셋째, 선택된 내부 표현이 엔지니어링 답변에 인과적으로 영향을 미친다. 이를 위해 Jacobian lens와 직접 읽기 기법을 결합하고, 60-law 반사적 벤치마크와 인과적 개입을 통해 내부 표현을 분석한다. 특히, 물리적 입력 방향만 반전된 동일한 프롬프트를 비교함으로써, 내부 상태 변화가 제시된 구성 법칙을 따르는지 확인한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM이 과학적 메커니즘을 내부적으로 표현하고 이를 제어할 수 있음을 보여주는 중요한 사례이다. 특히, Jacobian lens와 직접 읽기 기법을 통해 내부 표현을 분석하고, 인과적 개입을 통해 답변을 제어함으로써, LLM의 내부 작동 메커니즘을 더 깊이 이해할 수 있다. 그러나 일부 메커니즘은 직접 읽기보다 Jacobian lens가 효과적이지 않으며, 일부 프롬프트는 두 기법 모두에서 0의 변화를 보인다. 또한, 단어 집합만으로는 인과적 사용을 확정할 수 없으며, 일부 메커니즘은 수치적 비교와 동일하게 설명될 수 있다. 이는 LLM의 내부 표현이 항상 물리적 메커니즘을 반영하지는 않는다는 점을 시사한다.

실용적 활용

본 연구는 LLM의 내부 표현을 분석하고 제어하는 데 활용할 수 있으며, 특히 재료 과학 분야에서 모델의 신뢰성과 실패 진단에 기여할 수 있다. Jacobian lens와 인과적 개입 기법은 향후 강화 학습의 보상 신호 설계나 모델 훈련 목표 정의에 활용될 수 있으며, 과학적 결정이 맥락 변화에도 안정적으로 유지되는 AI 시스템 개발에 기여할 수 있다.