DocOps:面向复杂文档操作中自主智能体的可验证基准
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
July 22, 2026
作者: Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
cs.AI
摘要
随着自主智能体快速发展,其可靠操作普遍存在的数字文档的能力,已成为实现通用人工智能助手和自动化复杂工作空间工作流的关键。本文提出DocOps框架——一种确定性可验证的评估体系,该体系基于分层分类法,将源自真实实践的文档操作拆解为原子维度和递进式工作流复杂度。基于DocOps,我们系统评估了各类代理框架下具有代表性的闭源与开源模型,揭示出即使是最前沿的配置方案在处理高度耦合的长程任务时仍存在显著局限。进一步对现有智能体操作行为的细粒度分析,识别出三种关键故障模式:长期状态跟踪崩溃、浅层语义验证,以及结构元数据的破坏性编辑。最终,本研究揭示了智能体在维护全局文档一致性方面的能力边界,为未来在复杂数字生态系统中构建稳健、非破坏性的智能体指明了方向。
English
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.