Φ-Bench:大規模言語モデルは自らを支えるインフラストラクチャを設計・構築できるか?
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
September 9, 2026
著者: Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang
cs.AI
要旨
大規模言語モデル(LLM)は推論とコード生成において顕著な能力を示しており、自らを駆動する基盤そのものの開発と最適化を支援できる可能性が高まっている。しかし、既存のベンチマークは主に個別のカーネル、事前定義された演算子、あるいは事前指定された最適化対象に焦点を当てており、したがって、オープンエンドかつ長期にわたるLLMインフラストラクチャのエンジニアリングをLLMが遂行する能力を評価できていない。このギャップに対処するため、本稿では、LLMインフラストラクチャスタックのエンジニアリングに関してLLMを体系的に評価するためのベンチマークであるΦ-Benchを提示する。最先端研究で検討されている最適化問題に由来し、実世界のコードリポジトリに基づくΦ-Benchは、LLMインフラストラクチャスタックを広くカバーし、局所的なカーネルレベルの関数補完から長期にわたる実装およびエンドツーエンドのシステム最適化まで、複雑さの異なるタスクにまたがる。最先端LLMに対する広範な実験は、複雑なLLMインフラストラクチャのエンジニアリングにおける現在の能力と限界を明らかにし、将来のAIインフラストラクチャの自律的最適化へ向かう道筋に残る課題への洞察を与える。
English
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present Φ-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, Φ-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.