DocOps:針對複雜文檔操作中自主代理的可驗證基準
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
July 22, 2026
作者: Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
cs.AI
摘要
随着自主智能体的快速发展,其可靠操控各类数字文档的能力已成为实现通用人工智能助手及自动化复杂工作空间工作流的关键。本文提出DocOps——一个基于分层分类法的确定性可验证评估框架,该框架受真实实践启发,将文档操作解构为原子维度与逐步升级的工作流复杂性。基于DocOps,我们系统评估了代表性闭源与开源模型在不同智能体框架下的表现,揭示即使最先进的前沿配置在处理高耦合、长跨度任务时仍存在根本性局限。进一步,通过细粒度分析现有智能体的操控行为,我们识别出三种关键失败模式:长期状态追踪崩溃、浅层语义验证及结构性元数据的破坏性编辑。最终,本研究揭示了智能体在维护全局文档一致性方面的能力边界,为未来复杂数字生态系统中稳健、非破坏性智能体的设计提供了启示。
English
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.