ChatPaper.aiChatPaper

그래프 머신: 간선을 통한 더 나은 사전 학습을 향하여

Graph Machine: Towards Better Pretraining via Edges

September 2, 2026
저자: Lintai Hou
cs.AI

초록

우리는 O(n) 크기의 상태를 유지하고 이를 희소 동적 라우팅을 통해 접근하는 아키텍처인 그래프 머신(Graph Machine, GM)을 제안한다. 고정 크기 상태를 사용하거나 희소하지만 정적인 라우팅을 사용하는 방법들과 달리, GM은 잠재적으로 접근 가능한 상태 크기를 O(1)로 제한하지 않으면서 희소 계층에서 O(n) 복잡도를 유지한다. 대신 GM은 포인터 추적을 닮은 리퍼럴 메커니즘에 의해 미분 가능하게 갱신되는 포인터 유사 객체인 에지를 사용한다. 우리는 Qwen3-0.6B의 dense Transformer 계층 중 75%를 GM 희소 계층으로 교체하고 15.7B 토큰으로 처음부터 사전학습한다. 각 희소 계층에서 KV 헤드당 4,096개 토큰 중 단 2개만 검색할 때 손실은 약간만 저하되며, 4개일 때 최상의 모델은 손실을 소폭 개선한다.
English
We introduce the Graph Machine (GM), an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves O(n) complexity in its sparse layers without restricting the potentially accessible state size to O(1). Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best model marginally improves loss.