ChatPaper.aiChatPaper

DataSpace:面向异构工作空间的可验证分析数据代理基准测试

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

August 4, 2026
作者: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
cs.AI

摘要

数据代理使得在组织工作空间中进行自然语言分析成为可能,而相关证据可能分散在数据库、结构化文件、长文档和多模态媒体中。现有基准测试大多将结构化查询、检索或开放式分析彼此割裂,未能充分统一异构证据发现、完整表格输出以及确定性评估。我们提出DataSpace,这是一个基准测试,其中的数据代理从任务本地的异构工作空间中生成可验证的表格结果。它包含410个跨语言任务和7,439个工件,总计15.01 GB,涵盖CSV、JSON、SQLite、Markdown、PDF和视频格式。DataSpace还作为KDD Cup 2026复杂数据分析数据代理竞赛的官方评估基准。每个代理仅接收一个问题和工作空间,并返回所请求的完整表格结果。我们使用DataSpace-Builder构建DataSpace,这是一个基于执行的框架,包含跨语言转换、约束感知的关系采样、模态路由与工件渲染,以及由11位领域专家进行的人工审核和任务修复。确定性评估器执行表头无关的列对齐、类型和精度感知的归一化以及顺序感知的行比较。在六种最新发布的前沿多模态模型和五种广泛使用的代理框架中,最佳准确率达到66.34%,而在固定主干模型的情况下,框架选择造成15.36个百分点的差距。多模态证据整合与连接操作在所有六种主干模型上持续降低准确率。这些结果表明DataSpace尚未饱和,并揭示了提升数据代理可靠性的关键挑战。
English
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.