ChatPaper.aiChatPaper

Mechanist: 知性のメカニズムを発見するための科学機器としてのAI

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

August 12, 2026
著者: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
cs.AI

要旨

AIモデルは多様な領域で顕著な成功を収めているが、その能力の背後にあるメカニズムや、もたらし得るリスクについては、いまだ十分に理解されていない。AI開発の速度と自動化が進む一方で、機構の解明作業は依然として大部分が人手に依存しており、モデルの能力と、それを理解・制御する我々の能力との間の乖離は拡大し続けている。この乖離を埋めるため、我々はMechanistを提案する。Mechanistは、AIを科学的ツールとして活用し、AIの知能の背後にあるメカニズムを自律的に発見するエージェント型システムである。自律的な機構解明を支援するために、我々は解釈性に特化した約13,000件の論文からなる知識グラフを構築し、それを26分野にわたる4,300万件の論文からなる学際的なデータベースと統合した。さらに、メカニズム分析、因果介入、検証のための32の基盤的手法を厳選したライブラリを整備した。Claude Codeや既存のAI科学者システムと比較して、Mechanistはより価値の高いメカニズム仮説を生成し、実験をより確実に実行する。Mechanistはまた、モデルの挙動の発見から、AIモデルの説明と制御へと至る進展を示す。具体的には、まずMechanistは科学研究室における直感に反する安全性リスクを発見し、一見安全な訓練データを通じて安全でない特性がモダリティ間を越えて転移し得ることを明らかにする。次にMechanistは信念のメカニズム理論を構築し、モデルが世界知識をいかに表現するか、信念をいかに形成するか、他者の信念をいかに推論するか、そしてこれらのメカニズムが事前学習中にいかに出現するかを解明する。最後にMechanistは、これらの機構的洞察を実践的な介入策へと変換し、多様なシナリオにおけるモデルの性能を向上させるとともに、科学基盤モデルを指定された特性を持つDNA配列の生成へと導く。
English
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.