CausalDS: 데이터 과학 에이전트에서의 인과 추론 벤치마킹
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
July 9, 2026
저자: Andrej Leban, Yuekai Sun
cs.AI
초록
대규모 언어 모델(LLM)은 점점 더 통합 데이터 과학 에이전트 역할을 수행하며, 추상적 추론과 고급 도구 사용을 결합한다. 그러나 관련 벤치마크 환경은 대개 실제적인 데이터 분석이 없는 기호적 인과 추론 벤치마크와 원칙적인 인과 데이터 생성 구조가 없는 데이터 분석 벤치마크로 나뉜다. 더욱이, 기존의 인과 평가 데이터셋은 종종 기존 소스에서 선별된 예제에 국한되어 있으며, 다양성은 체계적인 새로운 합성 인과 구조 생성보다는 제한된 템플릿 변형에서 비롯된다. 우리는 에이전트 데이터 과학 워크플로우에서 인과 추론을 평가하기 위한 벤치마크인 CausalDS를 소개한다. 각 벤치마크 인스턴스는 샘플링된 구조적 인과 모델(SCM)과 생성된 관측 데이터, 그리고 현실적인 도메인에 기반한 합성 자연어 스토리로 구성된 장면(scene)이다. 우리는 선택적으로 벤치마크 구성 요소의 구성을 실제 데이터셋에서 얻은 경험적 분포에 기반하여, 완전히 합성된 생성을 통해 '인과 앵무새' 위험을 줄이면서 경험적 구조를 유지한다. 각 장면으로부터 우리는 Pearl의 세 단계를 모두 포괄하는 작업을 도출하며, 일반적인 데이터 과학 예측 작업은 1단계로 나타난다. 대부분의 작업에는 데이터 과학 코딩 구성 요소가 포함되며, 관측 모델에 의해 생성된 불완전한 관측이 자주 존재하기 때문에 모델은 일반적으로 여러 도구를 사용하여 최종 답변에 도달해야 한다. 또한, 질문에 정당한 답변이 없을 때 이를 인식하고 기권하는 것은 1등급 점수 결과로 처리된다. 따라서 이 벤치마크는 기호적 인과 추론, 데이터 과학, 불확실성 정량화, 기권, 도구 사용/코딩을 공동으로 평가한다.
English
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce CausalDS, a benchmark for evaluating causal reasoning in agentic data-science workflows. Each benchmark instance is a scene consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a realistic domain. We optionally ground the composition of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the "causal parrot" risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl's rungs, with typical data-science prediction tasks appearing as Rung 1. Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to the frequent presence of imperfect observations, which are generated by an observation model. Additionally, recognizing when a question admits no warranted answer and abstaining is treated as a first-class scored outcome. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.