Document Alignment based on Overlapping Fixed-Length Segments

Xiaotian Wang; Takehito Utsuro; Masaaki Nagata

doi:10.18653/v1/2024.acl-srw.10

Document Alignment based on Overlapping Fixed-Length Segments

Xiaotian Wang, Takehito Utsuro, Masaaki Nagata

Abstract

Acquiring large-scale parallel corpora is crucial for NLP tasks such asNeural Machine Translation, and web crawling has become a popularmethodology for this purpose. Previous studies have been conductedbased on sentence-based segmentation (SBS) when aligning documents invarious languages which are obtained through web crawling. Among them,the TK-PERT method (Thompson and Koehn, 2020) achieved state-of-the-artresults and addressed the boilerplate text in web crawling data wellthrough a down-weighting approach. However, there remains a problemwith how to handle long-text encoding better. Thus, we introduce thestrategy of Overlapping Fixed-Length Segmentation (OFLS) in place ofSBS, and observe a pronounced enhancement when performing the sameapproach for document alignment. In this paper, we compare the SBS andOFLS using three previous methods, Mean-Pool, TK-PERT (Thompson andKoehn, 2020), and Optimal Transport (Clark et al., 2019; El- Kishky andGuzman, 2020), on the WMT16 document alignment shared task forFrench-English, as well as on our self-established Japanese-Englishdataset MnRN. As a result, for the WMT16 task, various SBS basedmethods showed an increase in recall by 1% to 10% after reproductionwith OFLS. For MnRN data, OFLS demonstrated notable accuracyimprovements and exhibited faster document embedding speed.

Anthology ID:: 2024.acl-srw.10
Volume:: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Month:: August
Year:: 2024
Address:: Bangkok, Thailand
Editors:: Xiyan Fu, Eve Fleisig
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 51–61
Language:
URL:: https://aclanthology.org/2024.acl-srw.10
DOI:: 10.18653/v1/2024.acl-srw.10
Bibkey:
Cite (ACL):: Xiaotian Wang, Takehito Utsuro, and Masaaki Nagata. 2024. Document Alignment based on Overlapping Fixed-Length Segments. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 51–61, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):: Document Alignment based on Overlapping Fixed-Length Segments (Wang et al., ACL 2024)
Copy Citation:
PDF:: https://preview.aclanthology.org/dois-2013-emnlp/2024.acl-srw.10.pdf

PDF Search