NICT’s Submission to the WAT 2022 Structured Document Translation Task

Raj Dabre


Abstract
We present our submission to the structured document translation task organized by WAT 2022. In structured document translation, the key challenge is the handling of inline tags, which annotate text. Specifically, the text that is annotated by tags, should be translated in such a way that in the translation should contain the tags annotating the translation. This challenge is further compounded by the lack of training data containing sentence pairs with inline XML tag annotated content. However, to our surprise, we find that existing multilingual NMT systems are able to handle the translation of text annotated with XML tags without any explicit training on data containing said tags. Specifically, massively multilingual translation models like M2M-100 perform well despite not being explicitly trained to handle structured content. This direct translation approach is often either as good as if not better than the traditional approach of “remove tag, translate and re-inject tag” also known as the “detag-and-project” approach.
Anthology ID:
2022.wat-1.6
Volume:
Proceedings of the 9th Workshop on Asian Translation
Month:
October
Year:
2022
Address:
Gyeongju, Republic of Korea
Venue:
WAT
SIG:
Publisher:
International Conference on Computational Linguistics
Note:
Pages:
64–67
Language:
URL:
https://aclanthology.org/2022.wat-1.6
DOI:
Bibkey:
Cite (ACL):
Raj Dabre. 2022. NICT’s Submission to the WAT 2022 Structured Document Translation Task. In Proceedings of the 9th Workshop on Asian Translation, pages 64–67, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
Cite (Informal):
NICT’s Submission to the WAT 2022 Structured Document Translation Task (Dabre, WAT 2022)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-3/2022.wat-1.6.pdf