Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

Svetla Koeva; Nikola Obreshkov; Martin Yalamov

Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

Svetla Koeva, Nikola Obreshkov, Martin Yalamov

Abstract

The paper presents the Bulgarian MARCELL corpus, part of a recently developed multilingual corpus representing the national legislation in seven European countries and the NLP pipeline that turns the web crawled data into structured, linguistically annotated dataset. The Bulgarian data is web crawled, extracted from the original HTML format, filtered by document type, tokenised, sentence split, tagged and lemmatised with a fine-grained version of the Bulgarian Language Processing Chain, dependency parsed with NLP- Cube, annotated with named entities (persons, locations, organisations and others), noun phrases, IATE terms and EuroVoc descriptors. An orchestrator process has been developed to control the NLP pipeline performing an end-to-end data processing and annotation starting from the documents identification and ending in the generation of statistical reports. The Bulgarian MARCELL corpus consists of 25,283 documents (at the beginning of November 2019), which are classified into eleven types.

Anthology ID:: 2020.lrec-1.863
Volume:: Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:: May
Year:: 2020
Address:: Marseille, France
Editors:: Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:: LREC
SIG:
Publisher:: European Language Resources Association
Note:
Pages:: 6988–6994
Language:: English
URL:: https://preview.aclanthology.org/add-emnlp-2024-awards/2020.lrec-1.863/
DOI:
Bibkey:
Cite (ACL):: Svetla Koeva, Nikola Obreshkov, and Martin Yalamov. 2020. Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6988–6994, Marseille, France. European Language Resources Association.
Cite (Informal):: Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents (Koeva et al., LREC 2020)
Copy Citation:
PDF:: https://preview.aclanthology.org/add-emnlp-2024-awards/2020.lrec-1.863.pdf

PDF Cite Search Fix data