RST Continuity Corpus

Full Official Name: RST Continuity Corpus
Submission date: Oct. 15, 2024, 10:04 p.m.

RST Continuity Corpus was developed at Åbo Akademi University and Humboldt-Universität zu Berlin and contains annotations for continuity dimensions added to RST Discourse Treebank (LDC2002T07). RST Discourse Treebank is a collection of English news texts from the Penn Treebank annotated for rhetorical relations under the RST (Rhetorical Structure Theory) framework. In RST Continuity Corpus, the relations are annotated for the seven continuity dimensions: time, space, reference, action, perspective, modality, and speech act. The relations are also annotated for polarity, order of segments, nuclearity, and context. The source data consists of 1,009 relations from 217 Wall Street Journal texts annotated in RST Discourse Treebank for five relation types: causal, contrastive, conditional, elaboration and temporal. Annotation was performed using the UAM CorpusTool, version 2.8.16 or later. Files are presented as UTF-8 encoded XML and plain text. The corpus is divided into four sub-directories as described in the README file.

Creator(s)
Distributor(s)
Right Holder(s)