Mechanist:作为科学仪器探索智能机制的AI
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
August 12, 2026
作者: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
cs.AI
摘要
AI模型在众多领域取得了显著成功,然而其能力背后的机制及其可能带来的风险仍鲜为人知。随着AI开发速度日益加快且日趋自动化,机制性探索在很大程度上仍依赖人工操作,这导致模型的能力与我们理解和控制它们的能力之间的差距不断扩大。为弥合这一差距,我们引入了Mechanist——一个智能体系统,将AI作为科学仪器,用于自主发现AI智能背后的机制。为支持自主机制发现,我们构建了一个以可解释性为核心的知识图谱,涵盖约13,000篇论文,并将其与一个涵盖26个领域、包含4,300万篇论文的多学科数据库相整合。我们还精选了一个包含32种基础方法的工具库,用于机制分析、因果干预和验证。与Claude Code及现有的AI科学家系统相比,Mechanist能够产生更有价值的机制假设,并更可靠地执行实验。Mechanist还展示了从发现模型行为到解释和控制AI模型的递进过程。具体而言,Mechanist首先揭示了科学实验室中一个反直觉的安全风险,表明不安全特质可以通过表面安全的训练数据在不同模态之间转移。随后,Mechanist发展了一套信念的机制理论,揭示了模型如何表征世界知识、形成信念、推断他人信念,以及这些机制如何在预训练过程中涌现。最后,Mechanist将这些机制性洞见转化为实际干预措施,在多种场景下提升了模型性能,并引导科学基础模型生成具有指定属性的DNA序列。
English
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.