DataSpace: 이기종 워크스페이스에서 검증 가능한 분석을 위한 데이터 에이전트 벤치마킹
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
August 4, 2026
저자: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
cs.AI
초록
데이터 에이전트는 조직 작업 공간에서 자연어 기반 분석을 가능하게 하며, 관련 증거는 데이터베이스, 구조화된 파일, 장문 문서, 멀티미디어에 분산되어 있을 수 있다. 기존 벤치마크는 구조화된 질의, 검색, 또는 개방형 분석을 대부분 개별적으로 다루어, 이종 증거 발견, 완전한 표 형식 출력, 결정론적 평가를 충분히 통합하지 못하였다. 본 논문에서는 데이터 에이전트가 태스크별 이종 작업 공간에서 검증 가능한 표 형식 결과를 생성하는 벤치마크인 DataSpace를 소개한다. DataSpace는 410개의 교차 형식 태스크와 7,439개의 아티팩트를 포함하며, CSV, JSON, SQLite, Markdown, PDF, 비디오를 아우르는 총 15.01GB 규모이다. DataSpace는 또한 KDD Cup 2026 Data Agents for Complex Data Analysis 대회의 공식 평가 벤치마크로 활용되었다. 각 에이전트는 질문과 작업 공간만을 입력으로 받아 요청된 결과를 완전한 표 형식으로 반환한다. DataSpace는 실행 기반 프레임워크인 DataSpace-Builder를 통해 구축되었으며, 이는 교차 형식 변환, 제약 인식 관계형 샘플링, 모달리티 라우팅 및 아티팩트 렌더링, 그리고 11명의 도메인 전문가에 의한 인간 검토 및 태스크 수정을 포함한다. 결정론적 평가기는 헤더 불변 열 정렬, 유형 및 정밀도 인식 정규화, 순서 인식 행 비교를 수행한다. 최근 공개된 6개의 프론티어 멀티모달 모델과 널리 사용되는 5개의 에이전트 하네스를 평가한 결과, 최고 정확도는 66.34%에 도달했으며, 백본이 고정된 조건에서 하네스 선택에 따라 15.36포인트의 격차가 발생했다. 멀티모달 증거 통합과 조인은 6개 백본 모두에서 정확도를 일관되게 감소시켰다. 이러한 결과는 DataSpace가 아직 포화되지 않았음을 보여주며, 데이터 에이전트의 신뢰성 향상을 위한 핵심 과제를 식별한다.
English
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.