ChatPaper.aiChatPaper

Mechanist:以AI作為科學儀器揭示智能的機制

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

August 12, 2026
作者: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
cs.AI

摘要

AI模型在眾多領域取得了卓越成就,然而其能力背後的機制以及可能帶來的風險仍未被充分理解。隨著AI開發速度加快且日益自動化,機制探索在很大程度上仍依賴人工,這使得模型的能力與我們理解和控制它們的能力之間的差距不斷擴大。為彌補這一差距,我們提出了Mechanist——一個將AI作為科學儀器、用於自主發現AI智能背後機制的智慧體系統。為支持自主機制發現,我們建構了一個以可解釋性為核心、包含約13,000篇論文的知識圖譜,並將其與一個涵蓋26個領域、共4,300萬篇論文的跨學科資料庫進行整合。我們進一步整理了包含32種基礎方法的函式庫,涵蓋機制分析、因果干預與驗證。與Claude Code及現有的AI科學家系統相比,Mechanist能產生更有價值的機制假設,並更可靠地執行實驗。Mechanist亦展現了從發現模型行為到解釋與控制AI模型的進展。具體而言,Mechanist首先在科學實驗室中發現了一個違反直覺的安全風險,顯示不安全特徵可透過看似安全的訓練資料在模態之間轉移。接著,Mechanist發展了一套關於信念的機制理論,揭示了模型如何表徵世界知識、形成信念、推斷他人的信念,以及這些機制如何在預訓練過程中湧現。最後,Mechanist將這些機制洞見轉化為實際干預措施,在多種場景下提升模型效能,並引導科學基礎模型生成具有指定屬性的DNA序列。
English
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.