개방형 생성 평가를 위한 이중 수준 메타 루브릭: GAMUT, 사실적 완전성 벤치마크
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
July 21, 2026
저자: Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong
cs.AI
초록
긴 형식 생성물의 사실성 평가는 주로 정밀도, 즉 모델이 제시한 주장이 올바른지 측정하는 데 초점을 맞춰 왔다. 지배적인 분해-검색-검증 파이프라인은 부정확한 주장을 잘 포착하지만, 응답이 포함해야 할 모든 정보를 담고 있는지에 대해서는 거의 알려주지 않는다. 사실성의 누락된 절반인 사실적 완전성을 측정하는 것은 더 어렵다. 완전한 응답이 반드시 포함해야 할 사실들의 전체 집합을 열거해야 하기 때문인데, 이러한 사실들은 단순한 목록 형태를 이루는 경우가 드물다. 이들은 종종 포괄성이 중요한 개방형 집합, 순서가 있는 프로세스, 그리고 개별 이진 점검 목록으로는 포착하기 어려운 사실 간의 관계를 수반한다. 본 논문은 개방형 생성물 평가를 위한 2단계 메타 루브릭 프레임워크를 소개하고, 이를 긴 형식 생성물에서 사실적 완전성을 평가하는 벤치마크인 Gamut(근거 기반 다중 양식 사실성 평가)으로 구현한다. 이 프레임워크는 2단계 루브릭 표현에 기반한다. 구조화된 메타 루브릭이 요구되는 내용의 조직과 중요성을 포착하면, 이를 기계적으로 이진 형태의 기계 채점 가능한 평면 체크리스트 루브릭으로 컴파일하며, LLM 평가자가 이를 신뢰성 있게 채점한다. 우리는 10개 다양한 도메인에 걸쳐 실제 웨어러블 이미지에 기반한 1,813개의 질문을 구축했으며, 각 질문은 전문 인간 주석자가 검증한 증거 기반 루브릭과 쌍을 이룬다. 이 프레임워크는 양식에 구애받지 않으므로 텍스트 전용 변형도 함께 공개한다. 14개의 최첨단 및 오픈웨이트 모델을 평가한 결과, 이 벤치마크는 실질적으로 어렵고(Gemini 3.1 Pro의 최고 점수 58.7%), 변별력이 매우 높으며, 평가자 선택에 강건함을 확인했다.
English
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. They often involve open-ended sets where coverage is what matters, ordered processes, and relationships among facts that a list of independent boolean checks fails to capture. We introduce a two-level meta-rubric framework for evaluating open-ended generation, and instantiate it as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. The framework rests on a two-level rubric representation: a structured meta-rubric captures the organization and importance of the required content, which is then mechanically compiled into a flat checklist of binary, machine-gradable rubrics that an LLM judge scores reliably. We construct 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Because the framework is modality-agnostic, we also release a text-only variant. Evaluating 14 frontier and open-weight models, we find the benchmark genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.