FlavourBench:以可執行的烹飪真值對前沿語言模型進行排名
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
August 20, 2026
作者: Josef Chen, Erim Hayretci
cs.AI
摘要
開放式語言模型基準通常依賴繼承而來的評判者:人類偏好評審團、另一個模型,或脆弱的精確比對答案。我們提出 FlavourBench,一個自動化基準,其中一套具版本管理的烹飪系統提供密集且可執行的真實答案。每一項任務提供八種食材,並要求選出三種食材的組合;在模型執行之前,Epicure 已對全部 56 種可能的組合進行評分。我們在一個包含 534 項任務的相同核心上評估 27 個前沿模型端點,涵蓋替代、搭配與受限組合。每個排名模型在每個評審團與家族中恰好有 89 個有效回應(共 14,418 個模型-任務單元),從而消除排行榜中的差異性缺失。FlavourBench 分數為家族等權平均的凍結任務分數。我們使用 50,000 次錨定叢集拔靴複製來建立同時 95% 分數區間,並使用 100,000 次符號翻轉抽樣進行全部 351 組成對模型對比,且以 Holm 法控制多重比較。兩個獨立編制的評審團之間的相關係數為 r = 0.89(等級相關係數 rho = 0.80)。Grok 4.6 的點估計值最高,為 65.1(同時 95% 信賴區間 61.0–69.2);351 組模型配對中有 101 組獲得顯著區分。本次發布包含提示詞、所有組合分數對照表、原始回應、確切路徑、內容雜湊,以及一個可重現所有結果的離線驗證器。
English
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.