ChatPaper.aiChatPaper

DataSpace:異質工作空間中可驗證分析的資料代理器基準測試

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

August 4, 2026
作者: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
cs.AI

摘要

數據代理使組織工作空間上的自然語言分析成為可能,其中相關證據可能分散於資料庫、結構化檔案、長篇文件與多媒體之中。現有基準測試大多將結構化查詢、檢索或開放式分析相互分離,導致異質證據發現、完整表格輸出與確定性評估未能充分統合。我們引入DataSpace,這是一個要求數據代理從任務局部的異質工作空間中產出可驗證表格結果的基準測試。其包含410項跨語言任務與7,439個工件,共計15.01 GB,涵蓋CSV、JSON、SQLite、Markdown、PDF及影片等格式。DataSpace同時作為KDD Cup 2026「複雜數據分析之數據代理」競賽的官方評估基準。每個代理僅接收一個問題與一個工作空間,並須回傳完整的請求表格結果。我們以DataSpace-Builder建構DataSpace,這是一個以執行為根基的框架,包含跨語言轉換、約束感知的關聯式採樣、模態路由與工件渲染,以及由11位領域專家執行的人工審查與任務修復。確定性評估器執行表頭不變的欄位對齊、類型與精度感知的正規化,以及順序感知的列比對。在六個近期發布的前沿多模態模型與五個廣泛使用的代理框架中,最佳準確率達66.34%,而在固定骨幹模型的情況下,代理框架之選擇造成15.36個百分點的差異。多模態證據整合與聯接在所有六個骨幹模型上均持續降低準確率。這些結果顯示DataSpace尚未飽和,並揭示了提升數據代理可靠性的關鍵挑戰。
English
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.