Skip to content

Causality Detection

Causality detection is a sentence classification task: given a natural language sentence, does it express causal information?

The task can be framed as:

  • Binary classificationCausal (procausal or countercausal) vs. Uncausal
  • Ternary classificationProcausal / Countercausal / Uncausal

The library currently exposes the binary formulation via ClassLabel.

Data Schema

Each split is stored as a Parquet file with the following columns:

Column Type Description
index str Unique sentence identifier
text str The input sentence
label int ClassLabel.Causal (1) or ClassLabel.Uncausal (0)

What counts as causal?

A sentence is considered causal when it expresses a relation between two events A and B satisfying three conditions (following Grivaz):

  1. Temporal order — A precedes B (the effect cannot occur before the cause)
  2. Counterfactuality — B is less likely without A
  3. Ontological asymmetry — A causing B does not imply B causes A

Sentences that negate such a relation are countercausal and are still labelled Causal in the binary scheme. Common countercausal patterns include (Hagen et al. 2025)1:

Pattern Example
Direct negation "A does not cause B"
Lack of effect "A happened and B did not happen"
Inverse expected cause "B happened though A did not happen"
Usual inverse effect "B happened despite A"
Negated context "It is falsely believed that A causes B"
Violation of counterfactuality "A and B happened coincidentally"

Example

Input:  "The storm caused significant flooding."
Output: ClassLabel.Causal (1)

Input:  "She went to the store and bought milk."
Output: ClassLabel.Uncausal (0)

Input:  "Sugar does not cause hyperactivity."
Output: ClassLabel.Causal (1)   # countercausal

Datasets

Corpus Sentences Domain Year Reference Links
AltLex 1,000 Wikipedia 2016 (Hidey and McKeown 2016)2 📄 🤗
BECauSE 2.0 1,803 News 2017 (Dunietz et al. 2017)3 📄 🤗
BioCause 851 Medical 2013 (Mihaila et al. 2013)4 📄 🤗
Countercausal News Corpus (CCNC) 3,415 News 2025 (Hagen et al. 2025)1 📄
Causal News Corpus (CNC) 1,957 News 2022 (Tan et al. 2022)5 🤗
COPA 2,000 General 2011 (Roemmele et al. 2011)6 🤗
CausalTimeBank (CTB) 2,201 News 2014 (Mirza et al. 2014)7 📄 🤗
CaTeRS 488 Fiction 2016 (Mostafazadeh et al. 2016)8 📄 🤗
EventStoryLine (ESL) 2,247 News 2017 (Caselli and Vossen 2017)9 📄 🤗
EventCausality 583 Web 2011 (Do et al. 2011)10 🤗
FinCausal 2,136 Finance 2020
Penn Discourse Treebank 3.0 (PDTB 3.0) News / WSJ 2019 (Webber et al. 2019)11 🤗
PolitiCause 5,070 Politics 2024
SCITE 5,236 Science 2021 (Li et al. 2021)12 📄 🤗
SemEval-2010 Task 8 10,690 General 2010 (Hendrickx et al. 2010)13 🤗
TCR 172 News 2018 (Ning et al. 2018)14 📄 🤗
UniCausal 14,903 Multiple 2023 (Tan et al. 2023)15 📄

CCNC (Hagen et al. 2025)1 is the first dataset to explicitly distinguish procausal, countercausal, and uncausal sentences (inter-annotator agreement: Cohen's κ = 0.74).

Models

Models are evaluated using macro-averaged F₁.

Model F₁ Reference
DistilBERT 80.0% [@sanh2019distilbert]
RoBERTa 87.4% [@liu2019roberta]
Mistral-7B-Instruct 66.2% [@jiang2023mistral]

!!! note Models trained without countercausal examples misclassify countercausal sentences as causal more than 10× as often as models trained on CCNC (Hagen et al. 2025)1.


  1. Hagen, Tim, Niklas Deckers, Felix Wolter, Harrisen Scells, and Martin Potthast. 2025. “Investigating Counterclaims in Causality Extraction from Text.” CoRR abs/2510.08224. https://doi.org/10.48550/ARXIV.2510.08224

  2. Hidey, Christopher, and Kathy McKeown. 2016. “Identifying Causal Relations Using Parallel Wikipedia Articles.” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. https://doi.org/10.18653/V1/P16-1135

  3. Dunietz, Jesse, Lori S. Levin, and Jaime G. Carbonell. 2017. “The BECauSE Corpus 2.0: Annotating Causality and Overlapping Relations.” In Proceedings of the 11th Linguistic Annotation Workshop, LAW@EACL 2017, Valencia, Spain, April 3, 2017, edited by Nathan Schneider and Nianwen Xue. Association for Computational Linguistics. https://doi.org/10.18653/V1/W17-0812

  4. Mihaila, Claudiu, Tomoko Ohta, Sampo Pyysalo, and Sophia Ananiadou. 2013. “BioCause: Annotating and Analysing Causality in the Biomedical Domain.” BMC Bioinform. 14: 2. https://doi.org/10.1186/1471-2105-14-2

  5. Tan, Fiona Anting, Ali Hürriyetoglu, Tommaso Caselli, et al. 2022. “The Causal News Corpus: Annotating Causal Relations in Event Sentences from News.” In Proceedings of the Thirteenth Language Resources and Evaluation Conference, LREC 2022, Marseille, France, 20-25 June 2022, edited by Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, et al. European Language Resources Association. 

  6. Roemmele, Melissa, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. “Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning.” Logical Formalizations of Commonsense Reasoning, Papers from the 2011 AAAI Spring Symposium, Technical Report SS-11-06, Stanford, California, USA, March 21-23, 2011

  7. Mirza, Paramita, Rachele Sprugnoli, Sara Tonelli, and Manuela Speranza. 2014. “Annotating Causality in the TempEval-3 Corpus.” In Proceedings of the EACL 2014 Workshop on Computational Approaches to Causality in Language (CAtoCL), edited by Oleksandr Kolomiyets, Marie-Francine Moens, Martha Palmer, James Pustejovsky, and Steven Bethard. Association for Computational Linguistics. https://doi.org/10.3115/v1/W14-0702

  8. Mostafazadeh, Nasrin, Alyson Grealish, Nathanael Chambers, James F. Allen, and Lucy Vanderwende. 2016. “CaTeRS: Causal and Temporal Relation Scheme for Semantic Annotation of Event Structures.” In Proceedings of the Fourth Workshop on Events, EVENTS@HLT-NAACL 2016, San Diego, California, USA, June 17, 2016, edited by Martha Palmer, Eduard H. Hovy, Teruko Mitamura, and Tim O’Gorman. Association for Computational Linguistics. https://doi.org/10.18653/V1/W16-1007

  9. Caselli, Tommaso, and Piek Vossen. 2017. “The Event StoryLine Corpus: A New Benchmark for Causal and Temporal Relation Extraction.” In Proceedings of the Events and Stories in the News Workshop@ACL 2017, Vancouver, Canada, August 4, 2017, edited by Tommaso Caselli, Ben Miller, Marieke van Erp, et al. Association for Computational Linguistics. https://doi.org/10.18653/V1/W17-2711

  10. Do, Quang, Yee Seng Chan, and Dan Roth. 2011. “Minimally Supervised Event Causality Identification.” Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, EMNLP 2011, 27-31 July 2011, John McIntyre Conference Centre, Edinburgh, UK, A Meeting of SIGDAT, a Special Interest Group of the ACL, 294–303. 

  11. Webber, Bonnie, Rashmi Prasad, Alan Lee, and Aravind Joshi. 2019. “The Penn Discourse Treebank 3.0 Annotation Manual.” Philadelphia, University of Pennsylvania 35: 108. 

  12. Li, Zhaoning, Qi Li, Xiaotian Zou, and Jiangtao Ren. 2021. “Causality Extraction Based on Self-Attentive BiLSTM-CRF with Transferred Embeddings.” Neurocomputing 423: 207–19. https://doi.org/10.1016/J.NEUCOM.2020.08.078

  13. Hendrickx, Iris, Su Nam Kim, Zornitsa Kozareva, et al. 2010. “SemEval-2010 Task 8: Multi-Way Classification of Semantic Relations Between Pairs of Nominals.” In Proceedings of the 5th International Workshop on Semantic Evaluation, SemEval@ACL 2010, Uppsala University, Uppsala, Sweden, July 15-16, 2010, edited by Katrin Erk and Carlo Strapparava. The Association for Computer Linguistics. 

  14. Ning, Qiang, Zhili Feng, Hao Wu, and Dan Roth. 2018. “Joint Reasoning for Temporal and Causal Relations.” In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, edited by Iryna Gurevych and Yusuke Miyao. Association for Computational Linguistics. https://doi.org/10.18653/V1/P18-1212

  15. Tan, Fiona Anting, Xinyu Zuo, and See-Kiong Ng. 2023. “UniCausal: Unified Benchmark and Repository for Causal Text Mining.” In Big Data Analytics and Knowledge Discovery - 25th International Conference, DaWaK 2023, Penang, Malaysia, August 28-30, 2023, Proceedings, edited by Robert Wrembel, Johann Gamper, Gabriele Kotsis, A. Min Tjoa, and Ismail Khalil, vol. 14148. Lecture Notes in Computer Science. Springer. https://doi.org/10.1007/978-3-031-39831-5\23