Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs
Nikolay Paev, Stefan Marinov, Ivan Kratchanov, Petya Osenova, Kiril Simov
Abstract
The determine the reading order of the text extracted from a searchable PDF produced by an OCR software from an old newspaper is the first task in the process of preparation of corpora of old newspapers. In the paper we present an algorithm for generation of reading order of black selected from the corresponding PDF. Also we performed a tuning of the parameters of the algorithm. The optimization provides 10 % improvement.- Anthology ID:
- 2026.pressmint-1.10
- Volume:
- Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Editors:
- Maciej Ogrodniczuk, Petya Osenova, Tanja Wissik
- Venues:
- PressMint | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 56–64
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-pressmint-10
- DOI:
- 10.63317/336sd7ixv3fw
- Cite (ACL):
- Nikolay Paev, Stefan Marinov, Ivan Kratchanov, Petya Osenova, and Kiril Simov. 2026. Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs. In Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers, pages 56–64, Palma de Mallorca, Spain. Association for Computational Linguistics.
- Cite (Informal):
- Towards a Bulgarian Historical Newspaper Corpus – Construction of Reading Order over the Text in Searchable PDFs (Paev et al., PressMint 2026)