PDFMathTranslate: Scientific Document Translation Preserving Layouts

Rongxin Ouyang, Chang Chu, Zhikuang Xin, Xiangyao Ma


Abstract
Language barriers in scientific documents hinder the diffusion and development of science and technologies. However, prior efforts in translating such documents largely overlooked the information in layouts. To bridge the gap, we introduce PDFMathTranslate, the world’s first open-source software for translating scientific documents while preserving layouts. Leveraging the most recent advances in large language models and precise layout detection, we contribute to the community with key improvements in precision, flexibility, and efficiency. The work is open-sourced at https://github.com/byaidu/pdfmathtranslate with more than 222k downloads.
Anthology ID:
2025.emnlp-demos.71
Volume:
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
Month:
November
Year:
2025
Address:
Suzhou, China
Editors:
Ivan Habernal, Peter Schulam, Jörg Tiedemann
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
918–924
Language:
URL:
https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-demos.71/
DOI:
Bibkey:
Cite (ACL):
Rongxin Ouyang, Chang Chu, Zhikuang Xin, and Xiangyao Ma. 2025. PDFMathTranslate: Scientific Document Translation Preserving Layouts. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 918–924, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):
PDFMathTranslate: Scientific Document Translation Preserving Layouts (Ouyang et al., EMNLP 2025)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-demos.71.pdf