CausalDS: Benchmarking do Raciocínio Causal em Agentes de Ciência de Dados
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
July 9, 2026
Autores: Andrej Leban, Yuekai Sun
cs.AI
Resumo
Modelos de linguagem de grande porte (LLMs) atuam cada vez mais como agentes integrados de ciência de dados, combinando raciocínio abstrato com uso avançado de ferramentas. No entanto, o panorama relevante de benchmarks divide-se amplamente entre benchmarks de raciocínio causal simbólico sem análise realista de dados, ou benchmarks de análise de dados sem uma estrutura causal geradora de dados fundamentada. Além disso, os conjuntos de dados de avaliação causal existentes são frequentemente restritos a exemplos selecionados de fontes existentes, com diversidade proveniente de variações limitadas baseadas em modelos, em vez de geração sistemática de novas estruturas causais sintéticas. Apresentamos o CausalDS, um benchmark para avaliar o raciocínio causal em fluxos de trabalho agentivos de ciência de dados. Cada instância do benchmark é uma cena composta por um modelo causal estrutural (SCM) amostrado com dados observacionais gerados e uma história sintética em linguagem natural acompanhante, fundamentada em um domínio realista. Opcionalmente, fundamentamos a composição dos componentes do benchmark em distribuições empíricas obtidas de conjuntos de dados do mundo real, retendo assim a estrutura empírica enquanto reduzimos o risco de "papagaio causal" por meio de geração completamente sintética. A partir de cada cena, derivamos tarefas que abrangem todos os três degraus de Pearl, com tarefas típicas de previsão em ciência de dados aparecendo como Degrau 1. A maioria das tarefas inclui um componente de codificação em ciência de dados, no qual o modelo tipicamente precisa usar várias ferramentas para chegar à resposta final devido à presença frequente de observações imperfeitas, que são geradas por um modelo de observação. Além disso, reconhecer quando uma pergunta não admite uma resposta justificada e abster-se é tratado como um resultado pontuado de primeira classe. O benchmark, portanto, avalia conjuntamente raciocínio causal simbólico, ciência de dados, quantificação de incerteza, abstenção e uso de ferramentas/codificação.
English
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce CausalDS, a benchmark for evaluating causal reasoning in agentic data-science workflows. Each benchmark instance is a scene consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a realistic domain. We optionally ground the composition of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the "causal parrot" risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl's rungs, with typical data-science prediction tasks appearing as Rung 1. Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to the frequent presence of imperfect observations, which are generated by an observation model. Additionally, recognizing when a question admits no warranted answer and abstaining is treated as a first-class scored outcome. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.