ChatPaper.aiChatPaper

Φ-Bench:大型語言模型能否設計並建構支撐自身運作的基礎設施?

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

September 9, 2026
作者: Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang
cs.AI

摘要

大型語言模型(LLMs)已在推理與程式碼生成方面展現出卓越能力,使人們預期它們能協助開發與最佳化驅動其自身的基礎設施。然而,現有基準主要聚焦於孤立的核心(kernel)、預定義運算子或預先指定的最佳化目標,因此無法評估 LLMs 執行開放式、長時程 LLM 基礎設施工程的能力。為填補此一缺口,我們提出 Φ-Bench,一套用於系統性評估 LLMs 在 LLM 基礎設施堆疊工程上表現的基準。Φ-Bench 源自前沿研究中探討的最佳化問題,並奠基於真實世界的程式碼儲存庫,廣泛涵蓋 LLM 基礎設施堆疊,且橫跨各種複雜度的任務,從局部核心層級的函式補全,到長時程實作與端到端系統最佳化。對前沿 LLMs 的廣泛實驗揭示了其當前在工程複雜 LLM 基礎設施方面的能力與限制,並為未來 AI 基礎設施自主最佳化之路上仍存的挑戰提供洞見。
English
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present Φ-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, Φ-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.