ChatPaper.aiChatPaper

FlavourBench: 실행 가능한 요리 정답 데이터(ground truth)를 활용한 프론티어 언어 모델 순위 평가

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

August 20, 2026
저자: Josef Chen, Erim Hayretci
cs.AI

초록

개방형 언어모델 벤치마크는 일반적으로 인간 선호 패널, 다른 모델, 또는 취약한 정확 일치 키라는 평가자를 계승한다. 우리는 버전화된 요리 시스템이 조밀하고 실행 가능한 실측 자료를 제공하는 자동화된 벤치마크인 FlavourBench를 제안한다. 각 과제는 여덟 가지 재료를 제시하고 세 가지 재료로 구성된 포트폴리오를 요구하며, Epicure는 모델 실행 전에 가능한 56개 포트폴리오 전체를 채점한다. 우리는 대체, 페어링, 제약 조합을 포괄하는 동일한 534개 과제 코어에서 27개 프런티어 엔드포인트를 평가한다. 순위에 오른 모든 모델은 패널 및 계열별로 정확히 89개의 유효 응답을 가지므로(총 14,418개 모델-과제 셀), 리더보드에서 차등적 결측성이 제거된다. FlavourBench 점수는 고정된 과제 점수들의 동일 계열 평균이다. 우리는 동시 95% 점수 구간을 위해 50,000개의 앵커-클러스터 부트스트랩 복제를 사용하고, 홀름 제어 하에 총 351개 모델 쌍 대비를 위해 100,000개의 부호 반전 추출을 사용한다. 독립적으로 작성된 두 패널은 r = 0.89(순위 rho = 0.80)의 상관관계를 보인다. Grok 4.6이 65.1로 가장 큰 점추정치를 보이며(동시 95% 신뢰구간 61.0-69.2), 351개 모델 쌍 중 101개 쌍이 판정된다. 공개 자료에는 프롬프트, 모든 포트폴리오 점수 맵, 원시 응답, 정확한 경로, 콘텐츠 해시, 그리고 모든 결과를 재구성하는 오프라인 검증기가 포함된다.
English
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.