Mathilde Huguin
2026
Parallel Corpora of Scholarly Documents for English-French Machine Translation
Ziqian Peng | Lichao Zhu | Rachel Bawden | Maud Bénard | Éric de la Clergerie | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | François Yvon
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Ziqian Peng | Lichao Zhu | Rachel Bawden | Maud Bénard | Éric de la Clergerie | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | François Yvon
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
The growing ability of large language models (LLMs) to process long-range context opens new perspectives for document-level machine translation (MT), especially in scholarly communication. In fact, translating scholarly texts requires to integrate both local and long-range contextual information to ensure the consistency and coherence across the full document. However, document-level parallel corpora for such text types remain scarce, limiting both evaluation and domain adaptation of MT systems for this task. To address this gap, we introduce ParaEPS (Earth and Planetary Sciences Bilingual Corpus) and ParaNLP (Natural Language Processing Bilingual Corpus), two new parallel corpora covering 14k abstracts and 103 full-length articles in two scientific domains to be used for fine-tuning and evaluation purposes. We compare the performance of eight MT systems on these test sets and find that fine-tuning on document-level data closes the gap between open systems based on Large Language Models (LLMs) and commercial systems. We also find that the performance of recent LLMs can worsen when translating full articles instead of translating them on a per paragraph basisfine-tuning. These experiments underscore the need for corpora such as ParaEPS and ParaNLP.
2025
MaTOS: Machine Translation for Open Science
Rachel Bawden | Maud Bénard | Éric de la Clergerie | José Cornejo Cárcamo | Nicolas Dahan | Manon Delorme | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | Joachim Minder | Jean-François Nominé | Ziqian Peng | Laurent Romary | Panagiotis Tsolakis | Lichao Zhu | François Yvon
Proceedings of Machine Translation Summit XX: Volume 2
Rachel Bawden | Maud Bénard | Éric de la Clergerie | José Cornejo Cárcamo | Nicolas Dahan | Manon Delorme | Mathilde Huguin | Natalie Kübler | Paul Lerner | Alexandra Mestivier | Joachim Minder | Jean-François Nominé | Ziqian Peng | Laurent Romary | Panagiotis Tsolakis | Lichao Zhu | François Yvon
Proceedings of Machine Translation Summit XX: Volume 2
This paper is a short presentation of MaTOS, a project focusing on the automatic translation of scholarly documents. Its main aims are threefold: (a) to develop resources (term lists and corpora) for high-quality machine translation; (b) to study methods for translating complete, structured documents in a cohesive and consistent manner; (c) to propose novel metrics to evaluate machine translation in technical domains. Publications and resources are available on the project web site: https://anr-matos.gihub.io.
2024
Translate your Own: a Post-Editing Experiment in the NLP domain
Rachel Bawden | Ziqian Peng | Maud Bénard | Éric Clergerie | Raphaël Esamotunu | Mathilde Huguin | Natalie Kübler | Alexandra Mestivier | Mona Michelot | Laurent Romary | Lichao Zhu | François Yvon
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
Rachel Bawden | Ziqian Peng | Maud Bénard | Éric Clergerie | Raphaël Esamotunu | Mathilde Huguin | Natalie Kübler | Alexandra Mestivier | Mona Michelot | Laurent Romary | Lichao Zhu | François Yvon
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
The improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors.
2023
Le corpus « Machine Translation » : une exploration diachronique des (méta)données Istex
Mathilde Huguin | Sabine Barreaux
Actes de CORIA-TALN 2023. Actes de l'atelier "Analyse et Recherche de Textes Scientifiques" (ARTS)@TALN 2023
Mathilde Huguin | Sabine Barreaux
Actes de CORIA-TALN 2023. Actes de l'atelier "Analyse et Recherche de Textes Scientifiques" (ARTS)@TALN 2023
Le corpus Machine Translation se compose de publications scientifiques issues du réservoir Istex. Conçu comme un cas d’usage, il permet d’explorer l’histoire de la traduction automatique au travers des métadonnées et des textes intégraux disponibles pour chacun de ses documents. D’une part, les métadonnées permettent d’apporter un premier regard sur le paysage de la traduction automatique grâce à des tableaux de bord bibliométriques. D’autre part, l’utilisation d’outils de fouille de textes sur le texte intégral rend saillantes des informations inaccessibles sans une lecture approfondie des articles. L’exploration du corpus est réalisée grâce à Lodex, logiciel open source dédié à la valorisation de données structurées.
MaTOS: Traduction automatique pour la science ouverte
Maud Bénard | Alexandra Mestivier | Natalie Kubler | Lichao Zhu | Rachel Bawden | Eric De La Clergerie | Laurent Romary | Mathilde Huguin | Jean-François Nominé | Ziqian Peng | François Yvon
Actes de CORIA-TALN 2023. Actes de l'atelier "Analyse et Recherche de Textes Scientifiques" (ARTS)@TALN 2023
Maud Bénard | Alexandra Mestivier | Natalie Kubler | Lichao Zhu | Rachel Bawden | Eric De La Clergerie | Laurent Romary | Mathilde Huguin | Jean-François Nominé | Ziqian Peng | François Yvon
Actes de CORIA-TALN 2023. Actes de l'atelier "Analyse et Recherche de Textes Scientifiques" (ARTS)@TALN 2023
Cette contribution présente le projet MaTOS (Machine Translation for Open Science), qui vise à développer de nouvelles méthodes pour la traduction automatique (TA) intégrale de documents scientifiques entre le français et l’anglais, ainsi que des métriques automatiques pour évaluer la qualité des traductions produites. Pour ce faire, MaTOS s’intéresse (a) au recueil de ressources ouvertes pour la TA spécialisée; (b) à la description des marqueurs de cohérence textuelle pour les articles scientifiques; (c) au développement de nouvelles méthodes de traitement multilingue pour les documents; (d) aux métriques mesurant les progrès de la traduction de documents complets.