Adapting Multilingual Embedding Models to Historical Luxembourgish
Andrianos Michail, Corina Raclé, Juri Opitz, Simon Clematide
Abstract
The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical content due to OCR noise and outdated spellings. This study examines multilingual embeddings for cross-lingual semantic search in historical Luxembourgish (LB), a low-resource language. We collect historical Luxembourgish news articles from various periods and use GPT-4o for sentence segmentation and translation, generating 20,000 parallel training sentences per language pair. Additionally, we create a semantic search (Historical LB Bitext Mining) evaluation set and find that existing models perform poorly on cross-lingual search for historical Luxembourgish. Using our historical and additional modern parallel training data, we adapt several multilingual embedding models through contrastive learning or knowledge distillation and increase accuracy significantly for all models. We release our adapted models and historical Luxembourgish-German/French/English bitexts to support further research.- Anthology ID:
- 2025.latechclfl-1.26
- Volume:
- Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025)
- Month:
- May
- Year:
- 2025
- Address:
- Albuquerque, New Mexico
- Editors:
- Anna Kazantseva, Stan Szpakowicz, Stefania Degaetano-Ortlieb, Yuri Bizzoni, Janis Pagel
- Venues:
- LaTeCHCLfL | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 291–298
- Language:
- URL:
- https://preview.aclanthology.org/moar-dois/2025.latechclfl-1.26/
- DOI:
- 10.18653/v1/2025.latechclfl-1.26
- Cite (ACL):
- Andrianos Michail, Corina Raclé, Juri Opitz, and Simon Clematide. 2025. Adapting Multilingual Embedding Models to Historical Luxembourgish. In Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2025), pages 291–298, Albuquerque, New Mexico. Association for Computational Linguistics.
- Cite (Informal):
- Adapting Multilingual Embedding Models to Historical Luxembourgish (Michail et al., LaTeCHCLfL 2025)
- PDF:
- https://preview.aclanthology.org/moar-dois/2025.latechclfl-1.26.pdf