LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish

Nina Hosseini-Kivanani, Fred Philippy


Abstract
We present LuxBorrow, a borrowing-first analysis of Luxembourgish (LU) news spanning 27 years (1999–2025): 259,305 RTL articles and 43.7M tokens. Our pipeline combines sentence-level language identification (LU/DE/FR/EN) with a token-level borrowing resolver restricted to LU sentences, using lemmatization, a collected loanword registry, and compiled morphological/orthographic rules. Empirically, LU remains the matrix language across all documents, while multilingual practice is pervasive: 77.1% of articles include at least one donor language and 65.4% use three or four. Breadth does not imply intensity: median code-mixing index (CMI) increases from 3.90 (LU+1) to only 7.00 (LU+3), indicating localized insertions rather than balanced bilingual text. Domain/period summaries show moderate but persistent mixing, with CMI rising from 6.1 (1999–2007) to a peak of 8.4 (2020). Token-level adaptations total 25,444 instances and exhibit a mixed profile: morphological 63.8%, orthographic 35.9%, lexical 0.3%; the most frequent single rules are orthographic (on→oun, eur→er), while morphology is collectively dominant. Diachronically, code-switching intensifies, and morphologically adapted borrowings grow from a small base; French overwhelmingly supplies adapted items, with modest growth for German and negligible English. We advocate borrowing-centric evaluation, borrowed token/type rates, donor entropy over borrowed items, and assimilation ratios over headline document-level mixing indices.
Anthology ID:
2026.lrec-main.249
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3171–3183
Language:
URL:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.249/
DOI:
Bibkey:
Cite (ACL):
Nina Hosseini-Kivanani and Fred Philippy. 2026. LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish. International Conference on Language Resources and Evaluation, main:3171–3183.
Cite (Informal):
LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish (Hosseini-Kivanani & Philippy, LREC 2026)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.249.pdf