ChatPaper.aiChatPaper

オープンエンド生成評価のための二段階メタルーブリック:GAMUT(事実完全性のベンチマーク)

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

July 21, 2026
著者: Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong
cs.AI

要旨

長文生成の事実性評価は、主に精度(モデルが述べる主張が正しいかどうか)に焦点を当ててきた。主流である分解・検索・検証のパイプラインは、誤った主張をよく捉えるものの、応答が必要な情報をすべて含んでいるかどうかについてはほとんど示さない。事実完全性(事実性の欠けている半分)の測定はより困難である。なぜなら、完全な回答が含むべき事実の全集合を列挙する必要があり、これらの事実が単純な一覧になることは稀だからだ。それらは、カバレッジが重要となる非閉じた集合や、順序付けられたプロセス、独立したブールチェックのリストでは捉えられない事実間の関係を含むことが多い。本稿では、非閉じた生成を評価するための2段階メタルーブリックフレームワークを導入し、それを長文生成における事実完全性のベンチマークであるGamut(グラウンデッド・アセスメント・オブ・マルチモーダル・ファクチュアリティ)として具体化する。このフレームワークは、2段階のルーブリック表現に基づく。構造化されたメタルーブリックが、必要な内容の構成と重要度を捉え、それを機械的に、LLM判定器が信頼性高くスコアリングできる2値・機械評価可能なルーブリックのフラットなチェックリストにコンパイルする。我々は、10の多様な領域にわたる実際のウェアラブル画像に基づく1,813の質問を構築し、各質問には専門家による人間の注釈者によって検証されたエビデンスに基づくルーブリックを付与した。このフレームワークはモダリティに依存しないため、テキストのみのバリアントも公開する。14の最先端およびオープンウェイトモデルを評価した結果、本ベンチマークは真に挑戦的であり(最高スコアはGemini 3.1 Proの58.7%)、識別力が高く、判定器の選択に対してもロバストであることが判明した。
English
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. They often involve open-ended sets where coverage is what matters, ordered processes, and relationships among facts that a list of independent boolean checks fails to capture. We introduce a two-level meta-rubric framework for evaluating open-ended generation, and instantiate it as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. The framework rests on a two-level rubric representation: a structured meta-rubric captures the organization and importance of the required content, which is then mechanically compiled into a flat checklist of binary, machine-gradable rubrics that an LLM judge scores reliably. We construct 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Because the framework is modality-agnostic, we also release a text-only variant. Evaluating 14 frontier and open-weight models, we find the benchmark genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.