@inproceedings{mikelenic-tadic-2020-building,
    title = "Building the {S}panish-{C}roatian Parallel Corpus",
    author = "Mikeleni{\'c}, Bojana  and
      Tadi{\'c}, Marko",
    editor = "Calzolari, Nicoletta  and
      B{\'e}chet, Fr{\'e}d{\'e}ric  and
      Blache, Philippe  and
      Choukri, Khalid  and
      Cieri, Christopher  and
      Declerck, Thierry  and
      Goggi, Sara  and
      Isahara, Hitoshi  and
      Maegaard, Bente  and
      Mariani, Joseph  and
      Mazo, H{\'e}l{\`e}ne  and
      Moreno, Asuncion  and
      Odijk, Jan  and
      Piperidis, Stelios",
    booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference",
    month = may,
    year = "2020",
    address = "Marseille, France",
    publisher = "European Language Resources Association",
    url = "https://preview.aclanthology.org/ingest-emnlp/2020.lrec-1.484/",
    pages = "3932--3936",
    language = "eng",
    ISBN = "979-10-95546-34-4",
    abstract = "This paper describes the building of the first Spanish-Croatian unidirectional parallel corpus, which has been constructed at the Faculty of Humanities and Social Sciences of the University of Zagreb. The corpus is comprised of eleven Spanish novels and their translations to Croatian done by six different professional translators. All the texts were published between 1999 and 2012. The corpus has more than 2 Mw, with approximately 1 Mw for each language. It was automatically sentence segmented and aligned, as well as manually post-corrected, and contains 71,778 translation units. In order to protect the copyright and to make the corpus available under permissive CC-BY licence, the aligned translation units are shuffled. This limits the usability of the corpus for research of language units at sentence and lower language levels only. There are two versions of the corpus in TMX format that will be available for download through META-SHARE and CLARIN ERIC infrastructure. The former contains plain TMX, while the latter is lemmatised and POS-tagged and stored in the aTMX format."
}Markdown (Informal)
[Building the Spanish-Croatian Parallel Corpus](https://preview.aclanthology.org/ingest-emnlp/2020.lrec-1.484/) (Mikelenić & Tadić, LREC 2020)
ACL
- Bojana Mikelenić and Marko Tadić. 2020. Building the Spanish-Croatian Parallel Corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3932–3936, Marseille, France. European Language Resources Association.