ChatPaper.aiChatPaper

評估授權了什麼?一項關於Inspect Evals中主張相對推論的承諾界限普查

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

August 25, 2026
作者: Xi Qin
cs.AI

摘要

評估工件指定一個前向計算:任務、評分器與回報指標。它們不一定授權該指標所附帶的主張,因為重放該主張所需的歷史證據與替代語義可能未綁定。我們透過凍結基質 D、接地族 F、主張查詢 q 及由此產生的識別集合,將此缺失的主張重放層形式化。隨後,我們針對一固定提交版本,盤點全部 124 個機械上符合資格的 Inspect Evals 單元。每個單元皆獲得終止處置;其中 110 個在確定性推論之前即停止,因為所需的歷史證據或語義接地不可取得。在執行收斂之處,精確數值、勝者、完整排序與成對關係依主張解析及主要族與審查族之別而分離。因此,該審計回傳型別化停止、不穩定性見證與穩定子結構,而非強加單一的評估者意義或單一的穩健/非穩健標籤。
English
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.