Aligned Parallel Corpus of the Vedic Saṁhitās for Machine Translation

Yuzuki Tsukagoshi, Ikki Ohmukai


Abstract
We introduce a verse-/paragraph-aligned parallel corpus for three Vedic Saṁhitās –the R̥gveda (R̥V), the Atharvaveda Śaunaka (AVŚ), and the Taittirīya Saṁhitā (TS)– paired with authoritative public-domain translations (Geldner for R̥V, Whitney for AVŚ, and Keith for TS). The source texts are drawn from established digital editions (e.g., TITUS and VedaWeb) and normalized under ISO 15919. Each Sanskrit segment is aligned to exactly one translated unit (verse or paragraph for TS prose), yielding a unified, model-ready format. Using this resource, we fine-tune and evaluate three large language models –GPT-4.1 nano, Gemini 2.5 Flash, and Mitra– on VedicGerman/English translation. Evaluation combines surface and semantic metrics (case-insensitive sacreBLEU and COMET), enabling a balanced assessment of form and meaning. Results show consistent in-domain gains after supervised fine-tuning, but substantial cross-domain degradation when models are tested on unseen Saṁhitās, indicating pronounced stylistic and lexical divergence among R̥V, AVŚ, and TS. These findings motivate domain-aware training and reporting practices for Vedic machine translation. We release the corpus with standardized splits and preprocessing to support reproducibility and future d research on historical language modeling, alignment, and translation for low-resource ancient languages.
Anthology ID:
2026.lrec-main.272
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3434–3444
Language:
URL:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.272/
DOI:
Bibkey:
Cite (ACL):
Yuzuki Tsukagoshi and Ikki Ohmukai. 2026. Aligned Parallel Corpus of the Vedic Saṁhitās for Machine Translation. International Conference on Language Resources and Evaluation, main:3434–3444.
Cite (Informal):
Aligned Parallel Corpus of the Vedic Saṁhitās for Machine Translation (Tsukagoshi & Ohmukai, LREC 2026)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.272.pdf