ChatPaper.aiChatPaper

Mechanist: 지능의 메커니즘을 발견하기 위한 과학적 도구로서의 AI

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

August 12, 2026
저자: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
cs.AI

초록

AI 모델은 다양한 영역에서 놀라운 성과를 거두었지만, 그 능력의 기저에 있는 메커니즘과 잠재적 위험에 대한 이해는 여전히 부족하다. AI 개발이 더욱 빠르고 자동화됨에 따라, 메커니즘 탐구는 대부분 수작업에 의존하고 있어 모델이 할 수 있는 것과 이를 이해하고 통제하는 우리의 능력 사이의 격차는 점점 커지고 있다. 이러한 격차를 해소하기 위해 우리는 AI 지능의 기저 메커니즘을 자율적으로 발견하는 과학적 도구로 AI를 활용하는 에이전트 시스템인 Mechanist를 소개한다. 자율적 메커니즘 발견을 지원하기 위해 우리는 해석 가능성에 초점을 맞춘 약 13,000편의 논문으로 구성된 지식 그래프를 구축하고, 이를 26개 분야에 걸친 4,300만 편의 논문으로 구성된 다학제 데이터베이스와 통합했다. 또한 메커니즘 분석, 인과 개입, 검증을 위한 32가지 기초 방법론 라이브러리를 구축했다. Claude Code 및 기존 AI 과학자 시스템과 비교하여 Mechanist는 더 가치 있는 메커니즘 가설을 생성하고 실험을 더 안정적으로 수행한다. Mechanist는 또한 모델 행동 발견에서 AI 모델의 설명 및 제어로 이어지는 발전을 보여준다. 구체적으로 Mechanist는 먼저 과학 실험실에서 직관에 반하는 안전 위험을 발견하여, 겉보기에 안전한 훈련 데이터를 통해서도 불안전한 특성이 양식 간에 전이될 수 있음을 보여준다. 그런 다음 Mechanist는 믿음에 대한 메커니즘 이론을 개발하여 모델이 세계 지식을 어떻게 표현하고, 믿음을 형성하며, 타인의 믿음을 추론하고, 이러한 메커니즘이 사전 학습 중에 어떻게 출현하는지를 밝혀낸다. 마지막으로 Mechanist는 이러한 메커니즘 통찰력을 다양한 시나리오에서 모델 성능을 개선하고 과학 기반 모델이 특정 속성을 가진 DNA 서열을 생성하도록 유도하는 실용적 개입으로 전환한다.
English
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.