ChatPaper.aiChatPaper

DocOps: 複雑な文書操作における自律エージェントのための検証可能なベンチマーク

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

July 22, 2026
著者: Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
cs.AI

要旨

自律エージェントの急速な進化に伴い、汎用AIアシスタントの実現や複雑なワークスペースワークフローの自動化において、ユビキタスなデジタル文書を確実に操作する能力が極めて重要となっている。本論文では、実世界の実践から着想を得た文書操作を原子次元と段階的に複雑化するワークフローに分解する階層的タクソノミーに基づく、決定論的に検証可能な評価フレームワークDocOpsを提案する。DocOpsを用いて、様々なエージェンティックハーネス上で代表的なクローズドソース・オープンソースモデルを体系的に評価した結果、最先端のフロンティア構成でさえ、高度に結合された長距離タスクを処理する際に深刻な限界を示すことが明らかになった。さらに、既存エージェントの操作行動の詳細な分析により、長期的状態追跡の崩壊、浅い意味検証、構造的メタデータの破壊的編集という3つの主要な障害モードが特定された。最終的に、本研究はエージェントがグローバルな文書一貫性を維持する能力の限界を明らかにし、複雑なデジタルエコシステムにおける堅牢で非破壊的なエージェントの将来設計に示唆を与えるものである。
English
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.