ChatPaper.aiChatPaper

DataSpace: 異種ワークスペースにおける検証可能な分析のためのデータエージェントのベンチマーク評価

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

August 4, 2026
著者: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
cs.AI

要旨

データエージェントは、組織のワークスペース上での自然言語による分析を可能にする。そこでは、関連するエビデンスがデータベース、構造化ファイル、長文ドキュメント、マルチメディアに散在している可能性がある。既存のベンチマークは、構造化クエリ、検索、オープンエンド分析を概ね個別に扱っており、異種エビデンスの発見、完全な表形式出力、決定的評価を十分に統合していない。我々は、データエージェントがタスク固有の異種混在ワークスペースから検証可能な表形式結果を生成するベンチマークである DataSpace を提案する。DataSpace には、410 の言語横断タスクと 7,439 のアーティファクトが含まれ、CSV、JSON、SQLite、Markdown、PDF、ビデオにわたり合計 15.01 GB に及ぶ。DataSpace はまた、KDD Cup 2026 Data Agents for Complex Data Analysis コンペティションの公式評価ベンチマークとしても使用された。各エージェントは質問とワークスペースのみを受け取り、要求された完全な表形式結果を返す。我々は、DataSpace-Builder を用いて DataSpace を構築した。DataSpace-Builder は、言語横断変換、制約認識リレーショナルサンプリング、モダリティルーティングとアーティファクトレンダリング、および 11 名のドメイン専門家による人間レビューとタスク修正から構成される実行基盤型フレームワークである。決定的評価器は、ヘッダー不変の列整列、型および精度を考慮した正規化、順序認識の行比較を実行する。最近リリースされた 6 つの最先端マルチモーダルモデルと、広く使用されている 5 つのエージェントハーネスを用いた評価では、最高精度は 66.34% に達し、バックボーンを固定した場合のハーネス選択により 15.36 ポイントの差が生じた。マルチモーダルエビデンスの統合とジョインは、6 つのバックボーンすべてにおいて精度を一貫して低下させた。これらの結果は、DataSpace が未飽和であることを示し、データエージェントの信頼性向上に向けた重要な課題を明らかにする。
English
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.