Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen

arXiv:2608.12036 · 2026-08-13 공개 · arXiv · PDF

foundation-models knowledge-graph causal-intervention ai-scientist autonomous-discovery belief-theory mechanist ai-interpretability

Abstract

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.

한국어 요약

한 줄 요약

Mechanist는 AI 모델의 작동 메커니즘을 자동으로 탐색하는 시스템으로, 13,000개 논문 기반 지식그래프와 32개 메커니즘 해석 방법을 활용한다.

핵심 기여도

핵심 아이디어

AI 모델의 능력과 위험에 대한 메커니즘 이해는 여전히 수동적이다. Mechanist는 AI를 과학 도구로 활용하여 메커니즘을 자동으로 탐색하는 시스템이다. 핵심 아이디어는 해석 중심의 지식그래프와 32개 메커니즘 해석 방법을 통합하여, 가설 생성, 실험, 검증, 반복을 자동화하는 다단계 에이전트 프레임워크를 구축하는 것이다. 특히, Mechanist는 모델의 안전 위험을 탐지하고 믿음 형성 메커니즘을 이론화하며, 이를 실제 개입으로 전환한다.

기술적 접근법

주요 결과

의의 및 한계

Mechanist는 AI 모델의 메커니즘을 자동으로 탐색함으로써, 모델 이해와 제어의 격차를 줄이는 데 기여한다. 특히, 32개의 해석 방법과 13,000개 논문 기반 지식그래프는 고품질 가설 생성과 실험 실행을 가능하게 한다. 그러나, 반복 예산이 소진되면 미해결 문제와 한계와 함께 결과를 반환해야 하며, 이는 시스템의 완전한 자동화를 제한할 수 있다. 또한, 사용자의 초기 입력에 따라 가설의 방향성이 결정되므로, 입력의 질에 따라 결과가 달라질 수 있다.

실용적 활용

Mechanist는 AI 모델의 안전성 검증, 과학 실험실에서의 모델 활용, 그리고 DNA 시퀀스 생성과 같은 생명공학 분야에서 활용 가능하다. 또한, 모델의 믿음 형성 메커니즘을 이해함으로써, 인간과 AI 간의 신뢰 구축 및 협업을 촉진할 수 있다.