Tung Le
2025
Enhancing Scientific Title Generation via Optimized Sentence Ordering
Thien Khuu | Thuan Huynh | Nam Van Chi | Tung Le
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
Thien Khuu | Thuan Huynh | Nam Van Chi | Tung Le
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
HisGraphRAG: GraphRAG for Vietnamese Historical question answering
Hoang Thanh Nguyen | Tung Le | Huy Tien Nguyen
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
Hoang Thanh Nguyen | Tung Le | Huy Tien Nguyen
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
KWordinaryVQA: A Keyword-Driven Generative Visual Question Answering System for Culinary Exploration
Huy Trieu | Thanh Thai Nguyen | Thanh Nghia Vo | Thinh Vuong Vo | Thanh Tu Dang | Tung Le
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
Huy Trieu | Thanh Thai Nguyen | Thanh Nghia Vo | Thinh Vuong Vo | Thanh Tu Dang | Tung Le
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
KDA: Knowledge Distillation Adapter for Cross-Lingual Transfer
Ta-Bao Nguyen | Nguyen-Phuong Phan | Tung Le | Huy Tien Nguyen
Proceedings of the 18th International Natural Language Generation Conference
Ta-Bao Nguyen | Nguyen-Phuong Phan | Tung Le | Huy Tien Nguyen
Proceedings of the 18th International Natural Language Generation Conference
State-of-the-art cross-lingual transfer often relies on massive multilingual models, but their prohibitive size and computational cost limit their practicality for low-resource languages. An alternative is to adapt powerful, task-specialized monolingual models, but this presents challenges in bridging the vocabulary and structural gaps between languages. To address this, we propose KDA, a Knowledge Distillation Adapter framework that efficiently adapts a fine-tuned, high-resource monolingual model to a low-resource target language. KDA utilizes knowledge distillation to transfer the source model’s task-solving capabilities to the target language in a parameter-efficient manner. In addition, we introduce a novel adapter architecture that integrates source-language token embeddings while learning new positional embeddings, directly mitigating cross-lingual representational mismatches. Our empirical results on zero-shot transfer for Vietnamese Sentiment Analysis demonstrate that KDA significantly outperforms existing methods, offering a new, effective, and computationally efficient pathway for cross-lingual transfer.
2022
Bi-directional Cross-Attention Network on Vietnamese Visual Question Answering
Duy-Minh Nguyen-Tran | Tung Le | Minh Le Nguyen | Huy Tien Nguyen
Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation
Duy-Minh Nguyen-Tran | Tung Le | Minh Le Nguyen | Huy Tien Nguyen
Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation
VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering
Nguyen-Khang Le | Dieu-Hien Nguyen | Tung Le | Minh Nguyen
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Nguyen-Khang Le | Dieu-Hien Nguyen | Tung Le | Minh Nguyen
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Vietnamese is the native language of over 98 million people in the world. However, existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidence to explain the answer. We introduce VIMQA, a new Vietnamese dataset with over 10,000 Wikipedia-based multi-hop question-answer pairs. The dataset is human-generated and has four main features: (1) The questions require advanced reasoning over multiple paragraphs. (2) Sentence-level supporting facts are provided, enabling the QA model to reason and explain the answer. (3) The dataset offers various types of reasoning to test the model’s ability to reason and extract relevant proof. (4) The dataset is in Vietnamese, a low-resource language. We also conduct experiments on our dataset using state-of-the-art Multilingual single-hop and multi-hop QA methods. The results suggest that our dataset is challenging for existing methods, and there is room for improvement in Vietnamese QA systems. In addition, we propose a general process for data creation and publish a framework for creating multilingual multi-hop QA datasets. The dataset and framework are publicly available to encourage further research in Vietnamese QA systems.