Joungsu Choi
2025
DaCoM: Strategies to Construct Domain-specific Low-resource Language Machine Translation Dataset
Junghoon Kang
|
Keunjoo Tak
|
Joungsu Choi
|
Myunghyun Kim
|
Junyoung Jang
|
Youjin Kang
Proceedings of the 31st International Conference on Computational Linguistics: Industry Track
Translation of low-resource languages in industrial domains is essential for improving market productivity and ensuring foreign workers have better access to information. However, existing translators struggle with domain-specific terms, and there is a lack of expert annotators for dataset creation. In this work, we propose DaCoM, a methodology for collecting low-resource language pairs from industrial domains to address these challenges. DaCoM is a hybrid translation framework enabling effective data collection. The framework consists of a large language model and neural machine translation. Evaluation verifies existing models perform inadequately on DaCoM-created datasets, with up to 53.7 BLEURT points difference depending on domain inclusion. DaCoM is expected to address the lack of datasets for domain-specific low-resource languages by being easily pluggable into future state-of-the-art models and maintaining an industrial domain-agnostic approach.
Distilling Cross-Modal Knowledge into Domain-Specific Retrievers for Enhanced Industrial Document Understanding
Jinhyeong Lim
|
Jeongwan Shin
|
Seeun Lee
|
Seongdeok Kim
|
Joungsu Choi
|
Jongbae Kim
|
Chun Hwan Jung
|
Youjin Kang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Retrieval-Augmented Generation (RAG) has shown strong performance in open-domain tasks, but its effectiveness in industrial domains is limited by a lack of domain understanding and document structural elements (DSE) such as tables, figures, charts, and formula.To address this challenge, we propose an efficient knowledge distillation framework that transfers complementary knowledge from both Large Language Models (LLMs) and Vision-Language Models (VLMs) into a compact domain-specific retriever.Extensive experiments and analysis on real-world industrial datasets from shipbuilding and electrical equipment domains demonstrate that the proposed framework improves both domain understanding and visual-structural retrieval, outperforming larger baselines while requiring significantly less computational complexity.
Search
Fix author
Co-authors
- Youjin Kang 2
- Junyoung Jang 1
- Chun Hwan Jung 1
- Junghoon Kang 1
- Myunghyun Kim 1
- show all...