Resource: The ACL RD-TEC 2.0

Reference The ACL RD-TEC 2.0
Date of Submission March 8, 2016, 12:31 p.m.
Status accepted
ISLRN 178-746-003-351-0
Resource Type Manually Annotated Corpus of Terms in Context
Media Type Text
Source
Language English
Format/MIME Type TEXT/XML
Size 33216 tokens
Access Medium WWW
Description

The ACL Reference Dataset for Terminology Extraction and Classification version 2.0 (ACL RD-TEC 2.0) has been developed with the aim of providing a benchmark for the evaluation of methods for terminology extraction and classification as well as entity recognition tasks based on specialised text from the computational linguistics domain. This release of the corpus consists of 300 abstracts from articles in the ACL Anthology Reference Corpus, published between 1978--2006. In these abstracts, terms (i.e., single or multi-word lexical units with a specialised meaning) are manually annotated. In addition to their boundaries in running text, annotated terms are classified into one of the seven categories method, tool, language resource (LR), LR product, model, measures and measurements, and other. To assess the quality of the annotations and to determine the difficulty of this task, more than 171 of the abstracts are annotated twice, independently, by each of the two annotators. In total, 6,818 terms are identified and annotated, resulting in a specialised vocabulary made of 3,318 lexical forms, mapped to 3,471 concepts.

Version 2.0
Creator Behrang QasemiZadeh - DFG Collaborative Research Centre 991, Dusseldorf University , Anne-Kathrin Schumann - Saarland University
Distributor Behrang QasemiZadeh - DFG Collaborative Research Centre 991, Dusseldorf University , Anne-Kathrin Schumann - Saarland University
Rights Holder Behrang QasemiZadeh - DFG Collaborative Research Centre 991, Dusseldorf University , Anne-Kathrin Schumann - Saarland University