Causal Event Candidate Extraction
Given a sentence that has been classified as causal (or countercausal), causal event candidate extraction identifies the text spans that are plausible cause and effect candidates. The task is typically modelled as sequence labelling (BIO tagging) or as direct span prediction.
The output spans are candidates — their actual causal relationship is determined in the subsequent Causality Identification step.
Data Schema
Each split is stored as a Parquet file with the following columns:
| Column | Type | Description |
|---|---|---|
index |
str |
Unique sentence identifier |
text |
str |
The input sentence |
entity |
list[list[int, int]] |
Character-level [start, end] spans of candidate events |
Example
Input: "The storm caused significant flooding."
Output: entity = [[0, 9], [17, 36]]
# "The storm" → [0, 9]
# "significant flooding" → [17, 36]
Spans can overlap when a phrase participates in multiple causal pairs within the same sentence.
Datasets
| Corpus | Sentences | Domain | Year | Reference | Links |
|---|---|---|---|---|---|
| AltLex | 1,000 | Wikipedia | 2016 | (Hidey and McKeown 2016)1 | ![]() |
| BECauSE 2.0 | 1,803 | News | 2017 | (Dunietz et al. 2017)2 | ![]() |
| BioCause | 851 | Medical | 2013 | (Mihaila et al. 2013)3 | ![]() |
| Causal News Corpus (CNC) | 1,957 | News | 2022 | (Tan et al. 2022)4 | ![]() |
| COPA | 2,000 | General | 2011 | (Roemmele et al. 2011)5 | ![]() |
| CaTeRS | 488 | Fiction | 2016 | (Mostafazadeh et al. 2016)6 | ![]() |
| EventCausality | 583 | Web | 2011 | (Do et al. 2011)7 | ![]() |
| FinCausal | 2,136 | Finance | 2020 | — | — |
| Penn Discourse Treebank 3.0 (PDTB 3.0) | — | News / WSJ | 2019 | (Webber et al. 2019)8 | ![]() |
| SCITE | 5,236 | Science | 2021 | (Li et al. 2021)9 | ![]() |
| TCR | 172 | News | 2018 | (Ning et al. 2018)10 | ![]() |
| UniCausal | 14,903 | Multiple | 2023 | (Tan et al. 2023)11 | ![]() |
Models
Models are evaluated using macro-averaged F₁ over extracted spans.
| Model | F₁ | Reference |
|---|---|---|
| DistilBERT | 35.8% | [@sanh2019distilbert] |
| RoBERTa | 44.0% | [@liu2019roberta] |
!!! note Span extraction is substantially harder than sentence classification. The relatively low F₁ scores reflect the difficulty of localising exact event boundaries in free text.
-
Hidey, Christopher, and Kathy McKeown. 2016. “Identifying Causal Relations Using Parallel Wikipedia Articles.” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. https://doi.org/10.18653/V1/P16-1135. ↩
-
Dunietz, Jesse, Lori S. Levin, and Jaime G. Carbonell. 2017. “The BECauSE Corpus 2.0: Annotating Causality and Overlapping Relations.” In Proceedings of the 11th Linguistic Annotation Workshop, LAW@EACL 2017, Valencia, Spain, April 3, 2017, edited by Nathan Schneider and Nianwen Xue. Association for Computational Linguistics. https://doi.org/10.18653/V1/W17-0812. ↩
-
Mihaila, Claudiu, Tomoko Ohta, Sampo Pyysalo, and Sophia Ananiadou. 2013. “BioCause: Annotating and Analysing Causality in the Biomedical Domain.” BMC Bioinform. 14: 2. https://doi.org/10.1186/1471-2105-14-2. ↩
-
Tan, Fiona Anting, Ali Hürriyetoglu, Tommaso Caselli, et al. 2022. “The Causal News Corpus: Annotating Causal Relations in Event Sentences from News.” In Proceedings of the Thirteenth Language Resources and Evaluation Conference, LREC 2022, Marseille, France, 20-25 June 2022, edited by Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, et al. European Language Resources Association. ↩
-
Roemmele, Melissa, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. “Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning.” Logical Formalizations of Commonsense Reasoning, Papers from the 2011 AAAI Spring Symposium, Technical Report SS-11-06, Stanford, California, USA, March 21-23, 2011. ↩
-
Mostafazadeh, Nasrin, Alyson Grealish, Nathanael Chambers, James F. Allen, and Lucy Vanderwende. 2016. “CaTeRS: Causal and Temporal Relation Scheme for Semantic Annotation of Event Structures.” In Proceedings of the Fourth Workshop on Events, EVENTS@HLT-NAACL 2016, San Diego, California, USA, June 17, 2016, edited by Martha Palmer, Eduard H. Hovy, Teruko Mitamura, and Tim O’Gorman. Association for Computational Linguistics. https://doi.org/10.18653/V1/W16-1007. ↩
-
Do, Quang, Yee Seng Chan, and Dan Roth. 2011. “Minimally Supervised Event Causality Identification.” Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, EMNLP 2011, 27-31 July 2011, John McIntyre Conference Centre, Edinburgh, UK, A Meeting of SIGDAT, a Special Interest Group of the ACL, 294–303. ↩
-
Webber, Bonnie, Rashmi Prasad, Alan Lee, and Aravind Joshi. 2019. “The Penn Discourse Treebank 3.0 Annotation Manual.” Philadelphia, University of Pennsylvania 35: 108. ↩
-
Li, Zhaoning, Qi Li, Xiaotian Zou, and Jiangtao Ren. 2021. “Causality Extraction Based on Self-Attentive BiLSTM-CRF with Transferred Embeddings.” Neurocomputing 423: 207–19. https://doi.org/10.1016/J.NEUCOM.2020.08.078. ↩
-
Ning, Qiang, Zhili Feng, Hao Wu, and Dan Roth. 2018. “Joint Reasoning for Temporal and Causal Relations.” In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, edited by Iryna Gurevych and Yusuke Miyao. Association for Computational Linguistics. https://doi.org/10.18653/V1/P18-1212. ↩
-
Tan, Fiona Anting, Xinyu Zuo, and See-Kiong Ng. 2023. “UniCausal: Unified Benchmark and Repository for Causal Text Mining.” In Big Data Analytics and Knowledge Discovery - 25th International Conference, DaWaK 2023, Penang, Malaysia, August 28-30, 2023, Proceedings, edited by Robert Wrembel, Johann Gamper, Gabriele Kotsis, A. Min Tjoa, and Ismail Khalil, vol. 14148. Lecture Notes in Computer Science. Springer. https://doi.org/10.1007/978-3-031-39831-5\23. ↩

