Φ-Bench: 대규모 언어 모델은 자신들을 구동하는 인프라를 엔지니어링할 수 있는가?
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
September 9, 2026
저자: Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang
cs.AI
초록
대규모 언어 모델(LLM)은 추론과 코드 생성에서 놀라운 역량을 입증해 왔으며, 이는 이들이 자신을 구동하는 바로 그 인프라를 개발하고 최적화하는 데 기여할 수 있다는 전망을 제기한다. 그러나 기존 벤치마크는 주로 개별 커널, 사전 정의된 연산자 또는 사전 지정된 최적화 목표에 초점을 맞추고 있어, LLM이 개방형·장기적(long-horizon) LLM 인프라 엔지니어링을 수행하는 능력을 평가하지 못한다. 이러한 격차를 해결하기 위해, 우리는 LLM 인프라 스택 엔지니어링에 대한 LLM의 능력을 체계적으로 평가하기 위한 벤치마크인 Φ-Bench를 제안한다. 최첨단 연구에서 다루어지는 최적화 문제에서 파생되고 실제 코드 저장소에 기반한 Φ-Bench는 LLM 인프라 스택을 폭넓게 포괄하며, 국소적 커널 수준 함수 완성부터 장기적 구현 및 엔드투엔드 시스템 최적화에 이르기까지 다양한 복잡도의 작업을 아우른다. 최첨단 LLM에 대한 광범위한 실험은 복잡한 LLM 인프라를 엔지니어링하는 현재의 능력과 한계를 드러내며, 미래 AI 인프라의 자율 최적화를 향한 경로에 남아 있는 과제에 대한 통찰을 제공한다.
English
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present Φ-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, Φ-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.