UberText Corpus

We have collected and arranged a large (more than 6 Gb) amount of texts of Ukrainian periodicals and continue to extend our archive. We are aiming to reach one billion words.

The exact same texts are used for Word Embeddings calculation. Unfortunately, licence restrictions of some periodicals prohibit publishing their texts unchanged.

To give public access to this data, we split them into sentences and then shuffled them randomly. Thus, everyone can use these texts to calculate any statistical models that work on a sentence level. We also published the lemmatized version of these texts in a different archive. For tokenization and lemmatization, we used the nlp-uk library from Andriy Rysin and the BrUk group.

The archive contains sentences from the following periodicals:

Download

Корпус # of tokens # of sentences   Download Download the lemmatized version
News 461451019 31021650 1.1GB 951MB
Wikipedia 185645357 15786948 403MB 371MB
Fiction 18323509 1811548 41MB 38MB
Ubercorpus 665419885 48620146 1.6GB 1.5GB 

Information for licensors:

  • We distribute texts in a form of shuffled sentences, which makes it impossible to restore the original text.
  • Texts are distributed on condition of Fair Use for the needs of statistical and scientific analysis.
  • We are a non-profit organization and do not gain any advantage from distributing the mentioned materials.
  • If you have any comments on the texts distributed here, please contact us at [email protected]