@inproceedings{erjavec-2012-goo300k,
    title = "The goo300k corpus of historical {S}lovene",
    author = "Erjavec, Toma{\v{z}}",
    editor = "Calzolari, Nicoletta  and
      Choukri, Khalid  and
      Declerck, Thierry  and
      Do{\u{g}}an, Mehmet U{\u{g}}ur  and
      Maegaard, Bente  and
      Mariani, Joseph  and
      Moreno, Asuncion  and
      Odijk, Jan  and
      Piperidis, Stelios",
    booktitle = "Proceedings of the Eighth International Conference on Language Resources and Evaluation ({LREC}'12)",
    month = may,
    year = "2012",
    address = "Istanbul, Turkey",
    publisher = "European Language Resources Association (ELRA)",
    url = "https://preview.aclanthology.org/ingest-emnlp/L12-1232/",
    pages = "2257--2260",
    abstract = "The paper presents a gold-standard reference corpus of historical Slovene containing 1,000 sampled pages from over 80 texts, which were, for the most part, written between 1750-1900. Each page of the transcription has an associated facsimile and the words in the texts have been manually annotated with their modern-day equivalent, lemma and part-of-speech. The paper presents the structure of the text collection, the sampling procedure, annotation process and encoding of the corpus. The corpus is meant to facilitate HLT research and enable corpus based diachronic studies for historical Slovene. The corpus is encoded according to the Text Encoding Initiative Guidelines (TEI P5), is available via a concordancer and for download from \url{http://nl.ijs.si/imp/} under the Creative Commons Attribution licence."
}Markdown (Informal)
[The goo300k corpus of historical Slovene](https://preview.aclanthology.org/ingest-emnlp/L12-1232/) (Erjavec, LREC 2012)
ACL
- Tomaž Erjavec. 2012. The goo300k corpus of historical Slovene. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 2257–2260, Istanbul, Turkey. European Language Resources Association (ELRA).