ChatPaper.aiChatPaper

FlavourBench:実行可能な料理グラウンド・トゥルースを用いたフロンティア言語モデルのランキング

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

August 20, 2026
著者: Josef Chen, Erim Hayretci
cs.AI

要旨

オープンエンドの言語モデルベンチマークは通常、人間の嗜好パネル、別のモデル、または脆弱な完全一致キーといった判定者を継承する。本稿では、バージョン管理された料理システムが密で実行可能なグラウンドトゥルースを提供する自動ベンチマークであるFlavourBenchを紹介する。各タスクは8つの材料を提示し、3材料のポートフォリオを要求する。モデル実行に先立ち、Epicureは可能な56全てのポートフォリオをスコアリングする。我々は、代替、ペアリング、制約付き構成にわたる同一の534タスクのコア上で、27のフロンティアエンドポイントを評価する。ランク付けされた各モデルは、パネルおよびファミリーごとに正確に89件の有効な応答を有し(合計14,418モデル-タスクセル)、これによりリーダーボードから差別的欠落が排除される。FlavourBenchスコアは、固定タスクスコアのファミリー等重み平均である。同時95%スコア帯には50,000回のアンカークラスターブートストラップ反復を用い、Holm補正を伴う全351のペアモデル対比には100,000回の符号反転抽出を用いる。独立に作成された2つのパネルはr = 0.89(順位rho = 0.80)で相関する。Grok 4.6は最大の点推定値65.1を示し(同時95%信頼区間61.0-69.2)、351のモデルペアのうち101が確定された。リリースには、プロンプト、全てのポートフォリオスコアマップ、生の応答、正確なルート、コンテンツハッシュ、および全ての結果を再構成するオフライン検証器が含まれる。
English
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.