ChatPaper.aiChatPaper

用於評估開放式生成的雙層元評分規準:GAMUT,一個事實完整性基準

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

July 21, 2026
作者: Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong
cs.AI

摘要

評估長篇生成內容的事實性,主要聚焦於精確度,即衡量模型所提出的主張是否正確。主流的「分解-搜尋-驗證」流程確實能有效捕捉不正確的主張,但對於回應是否包含所有應有資訊,卻著墨甚少。衡量事實完整性——事實性的缺失一半——更加困難:它需要列舉一個完整答案應包含的全部事實,而這些事實很少能形成一份扁平清單。它們通常涉及開放的集合(其中覆蓋率才是關鍵)、有序的流程,以及事實間的關係,這些關係是獨立布林檢查清單所無法捕捉的。我們提出一個兩層級的元評分量表框架來評估開放式生成,並將其具體化為 Gamut(多模態事實性的基礎評估),一個針對長篇生成事實完整性的基準。該框架建立在兩層級的評分量表表示法之上:一個結構化的元評分量表捕捉所需內容的組織與重要性,然後機械性地編譯成一份由二元、機器可評分的評分量表組成的扁平檢查清單,讓大型語言模型評分者能可靠地打分。我們建構了 1,813 個問題,這些問題基於來自 10 個多元領域的真實可穿戴裝置影像,每個問題都附有由專家人工註釋者驗證的、有證據支持的評分量表。由於該框架與模態無關,我們也釋出了一個純文字版本。評估 14 個前沿模型與開放權重模型後,我們發現該基準確實具有挑戰性(最佳分數為 Gemini 3.1 Pro 的 58.7%),區分度極高,且對評分者的選擇具有穩健性。
English
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. They often involve open-ended sets where coverage is what matters, ordered processes, and relationships among facts that a list of independent boolean checks fails to capture. We introduce a two-level meta-rubric framework for evaluating open-ended generation, and instantiate it as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. The framework rests on a two-level rubric representation: a structured meta-rubric captures the organization and importance of the required content, which is then mechanically compiled into a flat checklist of binary, machine-gradable rubrics that an LLM judge scores reliably. We construct 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Because the framework is modality-agnostic, we also release a text-only variant. Evaluating 14 frontier and open-weight models, we find the benchmark genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.