ChatPaper.aiChatPaper

평가가 허가하는 것은 무엇인가? Inspect 평가에서 주장-상대적 추론의 커밋-바운드 총람

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

August 25, 2026
저자: Xi Qin
cs.AI

초록

평가 산출물(evaluation artifacts)은 작업(task), 스코어러(scorer), 보고된 지표(metric)라는 순방향 계산을 명세한다. 그러나 해당 지표에 부착된 주장이 반드시 정당화되는 것은 아니다. 그 주장을 재현하는 데 필요한 역사적 증거와 대안적 의미론이 결합되지 않을 수 있기 때문이다. 본 연구는 이러한 결여된 주장-재현 계층(claim-replay layer)을 고정된 기반 D(frozen substrate), 근거 기반 계열 F(grounded family), 주장 질의 q(claim query), 그리고 그 결과로 도출되는 식별 집합(identified set)을 통해 형식화한다. 이후 고정된 커밋(pinned commit) 시점의 기계적으로 적격한 124개의 Inspect Evals 단위를 전수 조사한다. 각 단위는 종착 상태(terminal disposition)를 부여받으며, 110개는 필수 역사적 증거나 의미적 근거 부여가 불가능하여 결정적 추론 이전에 중단된다. 실행이 종결되는 경우, 정확한 값, 승자, 완전한 순서, 쌍별 관계는 주장 해소(claim resolution) 및 1차(primary) 대 검토(review) 계열에 따라 구분된다. 따라서 본 감사는 단일 평가자 의미나 단일 견고/비견고 라벨을 강제하는 대신, 유형화된 중단(typed stops), 불안정성 증인(instability witnesses), 안정적 하위 구조(stable substructure)를 반환한다.
English
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.