Parsed Early English Books Online

Full Official Name: Parsed Early English Books Online - Text Creation Partnership
Submission date: Aug. 12, 2026, 7:49 p.m.

**Introduction** Parsed Early English Books Online - Text Creation Partnership (LDC2026T09) (EEBO-TCP) was developed by the Linguistic Data Consortium. It is a part-of-speech tagged and syntactically parsed version of the EEBO-TCP collection of Early Modern English texts. The corpus consists of 59,433 texts dating primarily from 1600-1700, with a smaller number of texts from earlier and later periods. Parsed EEBO-TCP was developed to support research in historical linguistics, the study of English syntax and language change, and natural language processing of historical texts. **Data** The corpus contains 48.6 million parsed sentences (trees) comprising more than 1.5 billion tokens. The parses were produced automatically using a parser trained on the Penn Parsed Corpus of Early Modern English, part of the Penn Parsed Corpora of Historical English (LDC2020T16). The automatically generated parses were not manually reviewed. Parsed EEBO-TCP includes the CorpusSearch 2 program and associated documentation. This tool allows users to search the data for syntactic structure, word sequences and words. An alternative version of CorpuSearch 2 for use on very large corproa is also included in this release. Text is provided in two formats: part-of-speech tagged text is presented as .pos files, and syntactically parsed text is presented as Penn Treebank-formatted .psd files. All text is UTF-8 encoded. **Updates** No updates at this time.

Creator(s)
Distributor(s)
Right Holder(s)