ChatPaper.aiChatPaper

評価は何を許可するのか?Inspect評価における主張相対的推論のコミット境界を考慮した調査

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

August 25, 2026
著者: Xi Qin
cs.AI

要旨

評価アーティファクトは、タスク、スコアラー、および報告メトリクスからなる順方向計算を指定する。しかし、それらは、そのメトリクスに付随する主張を必ずしも正当化するものではない。というのも、それを再現するために必要な歴史的証拠や代替的意味論が未結合のままである可能性があるからだ。我々は、この欠落した主張再現層を、凍結基盤D、接地された族F、主張クエリq、および結果として得られる識別集合によって形式化する。次に、固定されたコミット時点で機械的に適格な124のInspect Evalsユニットをすべて悉皆調査する。各ユニットは終局的な処遇を受ける。110は、必要な歴史的証拠または意味的接地が利用できないため、決定的推論の前に停止する。実行が完了する場合、正確な値、勝者、完全な順序、およびペアごとの関係は、主張の解決の有無と、一次ファミリーかレビューファミリーかによって区別される。したがって、監査は、単一の評価器の意味や単一の頑健/非頑健ラベルを強制するのではなく、型付き停止、不安定性の証人、および安定した部分構造を返す。
English
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.