模型還是支架?以互動為中心定位智能體故障之分類法
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
July 30, 2026
作者: Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru, Darvin Yi, Aakash Sabharwal, Yunzhong He
cs.AI
摘要
現有評估常將智能體失敗簡化為系統層級結果,掩蓋了錯誤的來源所在,以及何種干預措施能改善智能體系統。這形成了修復歸屬問題:同一可見失敗可能因來源不同而需要模型後訓練、框架工程、環境重新設計或基準修復。由於智能體行為源自模型、框架、使用者、工具、記憶與環境之間的互動,結果層級的標籤往往難以據以改善系統。多數失敗分類法因局限於特定基準且缺乏共享結構,對解決此問題幫助有限。我們提出一種以互動為中心的分類法,將失敗定位至其發源的互動,並識別應負責任的組成部分。該分類法將41種失敗模式組織起來,將每一種指派至兩個組成部分之間的邊,並標明修復歸屬所在的錯誤側。這使分類法具有可操作性:模型側失敗可識別後訓練目標,框架側失敗指向支架與工具整合的修復,而環境或評分器失敗則揭示需要重新設計的評估條件。此架構適用於各種智能體架構,從程式設計助手到長期任務個人助理以及多智能體系統。我們以公開基準、模型系統卡片、已發布報告及記錄的智能體軌跡中的具體實例為基礎建立此分類法,並使用獨立推理智能體作為評判者評估其可重現性。在四個前沿模型中,最強的評判者與人類類別標籤達到Cohen's κ=0.76,顯示這些類別捕捉到的是共享結構而非標註者特有的偏好。
English
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools, memory, and environments, outcome-level labels are often insufficient for improvement. Most failure taxonomies do little to resolve this problem because they are benchmark-specific and lack a shared structure. We introduce an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component. It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. This makes the taxonomy actionable: model-side failures identify targets for post-training, harness-side failures point to scaffolding and tool-integration fixes, and environment or grader failures reveal evaluation conditions requiring redesign. The schema applies across agent architectures, from coding assistants to long-horizon personal assistants and multi-agent systems. We ground the taxonomy in worked examples from public benchmarks, model system cards, published reports, and logged agent trajectories, and evaluate its reproducibility using independent reasoning agents as judges. Across four frontier models, the strongest judge reaches Cohen's κ=0.76 against human category labels, suggesting that the categories capture shared structure rather than annotator-specific preferences.