CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2609.11897v1Causal discovery asks a deceptively hard question: can we infer cause and effect from observed data alone? That matters in science, medicine, economics, and anywhere decisions depend on understanding interventions, not just correlations. The new benchmark CausalArena tackles a growing problem in this field: many methods look strong on one synthetic test, but their rankings change when the data-generating process changes. CausalArena unifies evaluation across several kinds of structural causal models, including synthetic graphs, semantically grounded scenarios, and formula-based scientific mechanisms, plus real-world datasets for an external check. The key idea is to test methods under a common protocol while varying the underlying causal families and assumptions. The result is sobering but useful: performance can shift dramatically across benchmark regimes, especially for foundation-model-based causal discovery. In other words, a method that wins one benchmark may not generalize to another. That makes CausalArena a valuable step toward more trustworthy causal AI.
Also spotted that day
Previous daily papers