ChatPaper.aiChatPaper

评估授予何种许可?——Inspect 评估中受提交约束的主张相对推断普查

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

August 25, 2026
作者: Xi Qin
cs.AI

摘要

评测工件规定了前向计算:任务、评分器和上报指标。它们并不必然授权附加于该指标之上的声明,因为重放该声明所需的历史证据和替代语义可能是未绑定的。我们通过冻结基底D、有据族F、声明查询q以及由此产生的识别集,将这一缺失的声明重放层形式化。随后,我们对固定提交版本下全部124个机制上合格的Inspect评测单元进行普查。每个单元均获得终态判定;其中110个在确定性推理之前即停止,因为所需的历史证据或语义依据不可用。在执行闭环之处,精确数值、胜者、完整排序和两两关系随声明解析及主族与审查族之分而分离。因此,该审计返回的是类型化停止、不稳定性证例和稳定子结构,而非强行赋予单一的评测者含义或单一的稳健/非稳健标签。
English
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.