Thìn Đặng Văn
Also published as: Dang Thin, Van Thin Dang, Dang Van Thin, Thin Dang Van, Thìn Đặng Văn, Dang Van Thin, Thin Dang Van
2026
DNT at #SMM4H–HeaRD 2026: Leveraging BERT-based Encoders and LLMs for Medical Information Extraction
Doan Nhat Tien | Thìn Đặng Văn
Proceedings of the 11th Social Media Mining for Health Research and Applications (SMM4H-HeaRD 2026) Workshop and Shared Tasks
Doan Nhat Tien | Thìn Đặng Văn
Proceedings of the 11th Social Media Mining for Health Research and Applications (SMM4H-HeaRD 2026) Workshop and Shared Tasks
This paper presents our systems for two tasks at #SMM4H-HeaRD 2026. For Task 1 (multilingual Adverse Drug Event detection), we fine-tune BERT-based multilingual models (InfoXLM and XLM-RoBERTa) and Qwen3.5-9B with ensemble methods, achieving 0.8584 macro F1 on the development set and 0.5304 F1 on unseen Farsi. For Task 7 (span detection of ClinicalImpacts and SocialImpacts in opioid narratives), DeBERTa-Large with simplified labeling achieves the best test performance (0.583 relaxed F1, 0.500 strict F1). Our analysis shows that LLMs excel on known languages in Task 1, while transformer-based models with simplified labeling generalize better for NER tasks.
CITD@UIT at SemEval-2026 Task 4: Structured Reasoning and Metric Specialization for Narrative Similarity
Thach Nguyen | Duc-Vu Nguyen | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Thach Nguyen | Duc-Vu Nguyen | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
We present a synergistic dual-track approach for SemEval-2026 Task 4 on narrative similarity, covering Track A (triple-wise classification) and Track B (narrative representation) through failure-driven data enrichment. The shared task received 71 final submissions from 46 teams across its two tracks. For Track A, we explore three reasoning strategies: hybrid Cross-Encoder–LLM arbitration (66.5% dev), DSPy-based component-wise decomposition (68.0% dev), and a multi-stage pairwise reasoning pipeline with enforced moral agency hierarchies, where the final Gemini 2.5 Pro/Flash system achieves 77.39% on development and 69.25% on test data, ranking 17th among 46 participating teams in the official evaluation. For Track B, we propose BGE-M3 (LoRA), an instruction-guided dense representation model trained with Multiple Negatives Ranking Loss (MNRL); since Track B provides only unlabeled story instances, we specialize the embedding space using adversarial samples synthesized from Track A failure cases, achieving 68.75% in the official evaluation and ranking 6th among 26 participating teams. Our analysis shows that narrative similarity depends more on outcome alignment and moral trajectory than lexical overlap, highlighting the complementary roles of explicit reasoning and task-specific metric-space specialization.
Stochastic Gradient Descenders at SemEval-2026 Task 9: Few-Shot LLM Prompting for Polarization Type Classification
Huynh Phu | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Huynh Phu | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
This paper presents our system for SemEval-2026 Task~9 (POLAR), Subtask~2, which focuses on classifying polarization types in social media text. We investigate three paradigms: (i) fine-tuning mDeBERTa-v3 with domain-adaptive pre-training, (ii) parameter-efficient adaptation of Qwen2.5-32B using LoRA, and (iii) few-shot prompting with Llama-3.3-70B-Instruct. Experimental results show that few-shot prompting, despite requiring no task-specific training, outperforms both fine-tuning and parameter-efficient approaches. Notably, it achieves non-zero F1 scores across all polarization categories, which is critical under macro-averaged evaluation. Our system ranks 2nd out of 29 English submissions on the official leaderboard, achieving an F1 Macro of 0.5157. These findings highlight the effectiveness of large instruction-tuned models in low-resource, label-imbalanced classification settings.
Gradient Descenders at SemEval-2026 Task 9: Data-Centric Counterfactual Augmentation for Multi-Label Hate Speech Detection
Tran Nhan | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Tran Nhan | Dang Thin
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
In this paper, we describe the Gradient Descenders submission to SemEval-2026 Task 9 Subtask 2: Multi-Label Hate Speech Detection. Existing Transformer-based approaches often exhibit degraded performance on this task due to severe class imbalance and complex class intersectionality, leading to the learning of spurious correlations. To counteract this, we introduce a novel, data-centric counterfactual augmentation pipeline. We employ Large Language Models (LLMs) as semantic generators to synthesize diverse, targeted training samples via three distinct prompting strategies: Additive Label-Flipping (Attribute Injection), Context Decoupling, and Cross-Domain Identity Substitution. Fine-tuning a RoBERTa classifier on this augmented corpus significantly improves the model’s sensitivity to minority classes. Ultimately, our system achieves a Macro-F1 score of 44.15% on the official test set, highlighting the efficacy of targeted LLM-based augmentation in highly imbalanced, multi-label environments.
An NLP Framework for Analyzing Corporate Strategic Behavior in the Opioid Industry Documents Archive
Duy Dang Phu | Thìn Đặng Văn
Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science
Duy Dang Phu | Thìn Đặng Văn
Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science
The Opioid Industry Documents Archive (OIDA) provides extensive internal corporate records that offer valuable insight into the drivers of the opioid crisis, yet its use in systematic analysis of corporate strategy remains limited. In this study, we propose an NLP-based framework to analyze strategic behavior in large-scale litigation archives, combining relevance filtering and topic modeling with large language model (LLM)-assisted interpretation. Applied to documents from Insys Therapeutics and Mallinckrodt Pharmaceuticals, our approach uncovers systematic differences in corporate strategies and organizational priorities. These results highlight the potential of integrating representation learning and LLMs for large-scale analysis in public health and corporate accountability research.
ViKhoMT: A Vietnamese–K’Ho Neural Machine Translation Dataset and Evaluation for Community Health Communication
Tram Truong | Vinh Nguyen | Dang Van Thin | Ngan Nguyen
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Tram Truong | Vinh Nguyen | Dang Van Thin | Ngan Nguyen
Proceedings of the Fifteenth Language Resources and Evaluation Conference
The Vietnamese government is prioritizing the socio-economic development and societal integration of ethnic minorities, including the K’Ho people. However, the lack of digital resources creates significant communication barriers, particularly in the critical domain of community health. To address this gap, we introduce ViKhoMT, a new, professionally curated Vietnamese-K’Ho parallel dataset containing approximately 10,000 sentence pairs focused on community health communication. To demonstrate the dataset’s quality and establish performance benchmarks, we conducted comprehensive evaluations by fine-tuning several pre-trained Neural Machine Translation (NMT) models. Our experiments show that a system based on the M2M100 architecture achieves BLEU scores of 60.5 for K’Ho-to-Vietnamese and 56.4 for Vietnamese-to-K’Ho, respectively. We release our dataset to the research community for free research purposes to support future studies and the development of practical translation tools for the K’Ho community. The dataset is publicly available at https://github.com/NgocTram2711/ViKhoMT.
EduPulse: A Practical LLM-Enhanced Opinion Mining System for Vietnamese Student Feedback in Educational Platforms
Xuan Phuc Nguyen | Xuan Phi Nguyen | Vinh Tiep Nguyen | Van Thin Dang | Ngan Luu-Thuy Nguyen
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track)
Xuan Phuc Nguyen | Xuan Phi Nguyen | Vinh Tiep Nguyen | Van Thin Dang | Ngan Luu-Thuy Nguyen
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track)
Opinion mining from real-world student feedback presents significant practical challenges, such as handling linguistic noise (slang, teencode) and the need for scalable and maintainable systems, which are often overlooked in academic research. This paper introduces EduPulse, a practical opinion mining system designed specifically to analyze student feedback in Vietnamese. Our application performs four opinion analysis tasks, including Sentiment Classification, Category-based Sentiment Classification, Suggestion Detection, and Opinion Summarization. We design the hybrid architecture that strategically balances performance, cost, and maintainability. This architecture leverages the robustness of Large Language Models (LLMs) for complex, noise-sensitive tasks as sentiment classification and suggestion detection, while employing a specialized, lightweight neural model for high-throughput, low-cost solutions. Our experiments show that applying the LLM-based approach achieves high robustness, justifying its operational cost by eliminating expensive retraining cycles. Furthermore, we demonstrate that our collaborative modular architecture significantly improves task performance (+7.6%) compared to traditional approaches, offering a practical design for industry-focused Natural Language Processing applications.
PhucNguyen@DravidianLangTech 2026: Political Multiclass Sentiment Analysis with XLM-RoBERTa and Low-Rank Adaptation
Dinh Khac Phuc Nguyen | Thìn Đặng Văn
Proceedings of the Sixth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages
Dinh Khac Phuc Nguyen | Thìn Đặng Văn
Proceedings of the Sixth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages
Analyzing political sentiment in code-mixed Tamil-English presents significant challenges due to informal jargon, severe class imbalance, and distribution shifts. This paper describes our system for the Political Multiclass Sentiment Analysis shared task at DravidianLangTech@ACL 2026, which categorizes tweets into seven sentiment classes. Our approach leverages XLM-RoBERTa integrated with Low-Rank Adaptation (LoRA). To mitigate majority-class dominance, we combine random oversampling with automated hyperparameter optimization to improve macro-level balance within this Parameter-Efficient Fine-Tuning (PEFT) framework. Enhanced by targeted preprocessing—specifically emoji demojization and noise removal—our system helps preserve nuanced symbolic cues, achieving a macro-average F1-score of 0.3763 and securing Rank 2 on the shared task leaderboard.
2025
Metamorphic at VLSP 2025: SIGMA – A Multimodal Agent System for Legal QA on Vietnamese Traffic Signs
Nguyen Tuan Kiet | Nguyen Khanh Tuan Anh | Long Hoang Huu Nguyen | Dam Vu Trong Tai | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Nguyen Tuan Kiet | Nguyen Khanh Tuan Anh | Long Hoang Huu Nguyen | Dam Vu Trong Tai | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Bosch@AI_Team at MMT 2025: Medical Machine Translation by Bidirectional Training with Small Language Models
Phan Minh Toan | Nguyen Xuan Phi | Nguyen Van Tai | Trang Minh Quang | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Phan Minh Toan | Nguyen Xuan Phi | Nguyen Van Tai | Trang Minh Quang | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
UIT-NTTT at VLSP2025: A Prompt Engineering Approach for Date Arithmetic Reasoning in Vietnamese
Khoa Nguyen-Anh Le | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Khoa Nguyen-Anh Le | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Bosch@AI_Team at LegalSML 2025: Vietnamese Legal Small Language with Domain Adaptation and Aspect-based Data Synthesis
Tran Minh Quang | Nguyen Xuan Phi | Nguyen Van Tai | Phan Minh Toan | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
Tran Minh Quang | Nguyen Xuan Phi | Nguyen Van Tai | Phan Minh Toan | Dang Van Thin
Proceedings of the 11th International Workshop on Vietnamese Language and Speech Processing
sonrobok4 Team at SemEval-2025 Task 8: Question Answering over Tabular Data Using Pandas and Large Language Models
Nguyen Minh Son | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Nguyen Minh Son | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
This paper describes the system of the son robok4 team for the SemEval-2025 Task 8: DataBench, Question-Answering over Tabular Data. The task requires answering questions based on the given question and dataset ID, ensuring that the responses are derived solely from the provided table. We address this task by using large language models (LLMs) to translate natural language questions into executable Python code for querying Pandas DataFrames. Furthermore, we employ techniques such as a rerun mechanism for error handling, structured metadata extraction, and dataset preprocessing to enhance performance. Our best-performing system achieved 89.46% accuracy on Subtask 1 and placed in the top 4 on the private test set. Additionally, it achieved 85.25% accuracy on Subtask 2 and placed in the top 9. We mainly focus on Subtask 1. We analyze the effectiveness of different LLMs for structured data reasoning and discuss key challenges in tabular question answering.
JellyK at SemEval-2025 Task 11: Russian Multi-label Emotion Detection with Pre-trained BERT-based Language Models
Khoa Anh-Nguyen Le | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Khoa Anh-Nguyen Le | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
This paper presents our approach for SemEval-2025 Task 11, we focus on on multi-label emotion detection in Russian text (track A). We preprocess the data by handling special characters, punctuation, and emotive expressions to improve feature-label relationships. To select the best model performance, we fine-tune various pre-trained language models specialized in Russian and evaluate them using K-FOLD Cross-Validation. Our results indicated that ruRoberta-large achieved the best Macro F1-score among tested models. Finally, our system achieved fifth place in the unofficial competition ranking.
NTA at SemEval-2025 Task 11: Enhanced Multilingual Textual Multi-label Emotion Detection via Integrated Augmentation Learning
Nguyen Pham Hoang Le | An Nguyen Tran Khuong | Tram Nguyen Thi Ngoc | Thin Dang Van
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Nguyen Pham Hoang Le | An Nguyen Tran Khuong | Tram Nguyen Thi Ngoc | Thin Dang Van
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Emotion detection in text is crucial for various applications, but progress, especially in multi-label scenarios, is often hampered by data scarcity, particularly for low-resource languages like Emakhuwa and Tigrinya. This lack of data limits model performance and generalizability. To address this, the NTA team developed a system for SemEval-2025 Task 11, leveraging data augmentation techniques: swap, deletion, oversampling, emotion-focused synonym insertion and synonym replacement to enhance baseline models for multilingual textual multi-label emotion detection. Our proposed system achieved significantly higher macro F1-scores compared to the baseline across multiple languages, demonstrating a robust approach to tackling data scarcity. This resulted in a 17th place overall ranking on the private leaderboard, and remarkably, we achieved the highest score and became the winner in Tigrinya language, demonstrating the effectiveness of our approach in a low-resource setting.
Firefly Team at SemEval-2025 Task 8: Question-Answering over Tabular Data using SQL/Python generation with Closed-Source Large Language Models
Ho Thuy Nga | Ho Thi Thanh Tuyen | Le Minh Hung | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Ho Thuy Nga | Ho Thi Thanh Tuyen | Le Minh Hung | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
In this paper, we describe our official system of the Firefly team for two main tasks in the SemEval-2025 Task 8: Question-Answering over Tabular Data. Our solution employs large language models (LLMs) to translate natural language queries into executable code, specifically Python and SQL, which are then used to generate answers categorized into five predefined types. Our empirical evaluation highlights the superiority of Python code generation over SQL for this challenge. Besides, the experimental results show that our system has achieved competitive performance in two subtasks. In Subtask I: Databench QA, where we rank the Top 9 across datasets of any size. Besides, our solution achieved competitive results and ranked 5th place in Subtask II: Databench QA Lite, where datasets are restricted to a maximum of 20 rows.
ABCD at SemEval-2025 Task 9: BERT-based and Generation-based models combine with advanced weighted majority soft voting strategy
Le Duc Tai | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Le Duc Tai | Dang Van Thin
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
This paper illustrates our ABCD team system approach in ACL 2025 - SemEval-2025 Task 9: The Food Hazard Detection Challenge, aim to solving both Task 1: Text classification for food hazard prediction, predicting the type of hazard and product, and Task 2: Food hazard and product “vector” detection, predicting the exact hazard and product. Precisely, we received a food report and our system needed to automatically detect which category of hazard and product the food belonged to. However, in Task 2, we must classify the food report into the exact name of the food hazard and category. To tackle Task 1, we implement and investigate various solutions, including (1) experimenting with a large battery of BERT-based models; and (2) utilizing generation-based models, and (3) taking advantage of a custom ensemble learning method. In addition, to address Task 2, we make use of different data augmentation techniques like synonym replacement and back-translation. To enhance the context of input, we cleaned some special characters that bring more clarity into text input. Our best official results on Task 1 and Task 2 are 0.786 and 0.458 in terms of F1-score, respectively—finally, our team solution achieved top 8th in task 1 and top 10th in task 2.
A.M.P at SciHal2025: Automated Hallucination Detection in Scientific Content via LLMs and Prompt Engineering
Le Nguyen Anh Khoa | Thìn Đặng Văn
Proceedings of the Fifth Workshop on Scholarly Document Processing (SDP 2025)
Le Nguyen Anh Khoa | Thìn Đặng Văn
Proceedings of the Fifth Workshop on Scholarly Document Processing (SDP 2025)
This paper presents our system developed for SciHal2025: Hallucination Detection for Scientific Content. The primary goal of this task is to detect hallucinated claims based on the corresponding reference. Our methodology leverages strategic prompt engineering to enhance LLMs’ ability to accurately distinguish between factual assertions and hallucinations in scientific contexts. Moreover, we discovered that aggregating the fine-grained classification results from the more complex subtask (subtask 2) into the simplified label set required for the simpler subtask (subtask 1) significantly improved performance compared to direct classification for subtask 1. This work contributes to the development of more reliable AI-powered research tools by providing a systematic framework for hallucination detection in scientific content.
twinhter at LeWiDi-2025: Integrating Annotator Perspectives into BERT for Learning with Disagreements
Nguyen Huu Dang Nguyen | Dang Van Thin
Proceedings of the The 4th Workshop on Perspectivist Approaches to NLP
Nguyen Huu Dang Nguyen | Dang Van Thin
Proceedings of the The 4th Workshop on Perspectivist Approaches to NLP
Annotator-provided information during labeling can reflect differences in how texts are understood and interpreted, though such variation may also arise from inconsistencies or errors. To make use of this information, we build a BERT-based model that integrates annotator perspectives and evaluate it on four datasets from the third edition of the Learning With Disagreements (LeWiDi) shared task. For each original data point, we create a new (text, annotator) pair, optionally modifying the text to reflect the annotator’s perspective when additional information is available. The text and annotator features are embedded separately and concatenated before classification, enabling the model to capture individual interpretations of the same input. Our model achieves first place on both tasks for the Par and VariErrNLI datasets. More broadly, it performs very well on datasets where annotators provide rich information and the number of annotators is relatively small, while still maintaining competitive results on datasets with limited annotator information and a larger annotator pool.
Exploring the Power of Large Language Models for Vietnamese Implitcit Sentiment Analysis
Huy Gia Luu | Dang Van Thin
Proceedings of the 18th International Natural Language Generation Conference
Huy Gia Luu | Dang Van Thin
Proceedings of the 18th International Natural Language Generation Conference
We present the first benchmark for implicit sentiment analysis (ISA) in Vietnamese, aimed at evaluating large language models (LLMs) on their ability to interpret implicit sentiment accompanied by ViISA, a dataset specifically constructed for this task. We assess a variety of open-source and close-source LLMs using state-of-the-art (SOTA) prompting techniques. While LLMs achieve strong recall, they often misclassify implicit cues such as sarcasm and exaggeration, resulting in low precision. Through detailed error analysis, we highlight key challenges and suggest improvements to Chain-of-Thought prompting via more contextually aligned demonstrations.
Automotive Document Labeling Using Large Language Models
Dang Van Thin | Cuong Xuan Chu | Christian Graf | Tobias Kaminski | Trung-Kien Tran
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Dang Van Thin | Cuong Xuan Chu | Christian Graf | Tobias Kaminski | Trung-Kien Tran
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Repairing and maintaining car parts are crucial tasks in the automotive industry, requiring a mechanic to have all relevant technical documents available. However, retrieving the right documents from a huge database heavily depends on domain expertise and is time consuming and error-prone. By labeling available documents according to the components they relate to, concise and accurate information can be retrieved efficiently. However, this is a challenging task as the relevance of a document to a particular component strongly depends on the context and the expertise of the domain specialist. Moreover, component terminology varies widely between different manufacturers. We address these challenges by utilizing Large Language Models (LLMs) to enrich and unify a component database via web mining, extracting relevant keywords, and leveraging hybrid search and LLM-based re-ranking to select the most relevant component for a document. We systematically evaluate our method using various LLMs on an expert-annotated dataset and demonstrate that it outperforms the baselines, which rely solely on LLM prompting.
Few-Shot Coreference Resolution with Semantic Difficulty Metrics and In-Context Learning
Nguyen Xuan Phuc | Dang Van Thin
Proceedings of the Eighth Workshop on Computational Models of Reference, Anaphora and Coreference
Nguyen Xuan Phuc | Dang Van Thin
Proceedings of the Eighth Workshop on Computational Models of Reference, Anaphora and Coreference
This paper presents our submission to the CRAC 2025 Shared Task on Multilingual Coreference Resolution in the LLM track. We propose a prompt-based few-shot coreference resolution system where the final inference is performed by Grok-3 using in-context learning. The core of our methodology is a difficulty- aware sample selection pipeline that leverages Gemini Flash 2.0 to compute semantic diffi- culty metrics, including mention dissimilarity and pronoun ambiguity. By identifying and selecting the most challenging training sam- ples for each language, we construct highly informative prompts to guide Grok-3 in predict- ing coreference chains and reconstructing zero anaphora. Our approach secured 3rd place in the CRAC 2025 shared task.
MMLabUIT at CoMeDiShared Task: Text Embedding Techniques versus Generation-Based NLI for Median Judgment Classification
Tai Duc Le | Thin Dang Van
Proceedings of Context and Meaning: Navigating Disagreements in NLP Annotation
Tai Duc Le | Thin Dang Van
Proceedings of Context and Meaning: Navigating Disagreements in NLP Annotation
This paper presents our approach in the COLING2025-CoMeDi task in 7 languages, focusing on sub-task 1: Median Judgment Classification with Ordinal Word-in-Context Judgments (OGWiC). Specifically, we need to determine the meaning relation of one word in two different contexts and classify the input into 4 labels. To address sub-task 1, we implement and investigate various solutions, including (1) Stacking, Averaged Embedding techniques with a multilingual BERT-based model; and (2) utilizing a Natural Language Inference approach instead of a regular classification process. All the experiments were conducted on the P100 GPU from the Kaggle platform. To enhance the context of input, we perform Improve Known Data Rate and Text Expansion in some languages. For model focusing purposes Custom Token was used in the data processing pipeline. Our best official results on the test set are 0.515, 0.518, and 0.524 in terms of Krippendorff’s α score on task 1. Our participation system achieved a Top 3 ranking in task 1. Besides the official result, our best approach also achieved 0.596 regarding Krippendorff’s α score on Task 1.
Baoflowin502 at MAHED Shared Task: Text-based Hate and Hope Speech Classification
Nguyen Minh Bao | Dang Van Thin
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Nguyen Minh Bao | Dang Van Thin
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
TranTranUIT at MAHED Shared Task: Multilingual Transformer Ensemble with Advanced Data Augmentation and Optuna-based Hyperparameter Optimization
Trinh Tran Tran | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Trinh Tran Tran | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
LoveHeaven at MAHED 2025: Text-based Hate and Hope Speech Classification Using AraBERT-Twitter Ensemble
Nguyễn Thiên Bảo | Dang Van Thin
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Nguyễn Thiên Bảo | Dang Van Thin
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
NguyenTriet at MAHED Shared Task: Ensemble of Arabic BERT Models with Hierarchical Prediction and Soft Voting for Text-Based Hope and Hate Detection
Nguyen Minh Triet | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Nguyen Minh Triet | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
912 at TAQEEM 2025: A Distribution-aware Approach to Arabic Essay Scoring
Trong-Tai Dam Vu | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Trong-Tai Dam Vu | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
PuxAI at QIAS 2025: Multi-Agent Retrieval-Augmented Generation for Islamic Inheritance and Knowledge Reasoning
Nguyen Xuan Phuc | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Nguyen Xuan Phuc | Thìn Đặng Văn
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
2024
NRK at SemEval-2024 Task 1: Semantic Textual Relatedness through Domain Adaptation and Ensemble Learning on BERT-based models
Nguyen Tuan Kiet | Dang Van Thin
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
Nguyen Tuan Kiet | Dang Van Thin
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
This paper describes the system of the team NRK for Task A in the SemEval-2024 Task 1: Semantic Textual Relatedness (STR). We focus on exploring the performance of ensemble architectures based on the voting technique and different pre-trained transformer-based language models, including the multilingual and monolingual BERTology models. The experimental results show that our system has achieved competitive performance in some languages in Track A: Supervised, where our submissions rank in the Top 3 and Top 4 for Algerian Arabic and Amharic languages. Our source code is released on the GitHub site.
Prompt Engineering with Large Language Models for Vietnamese Sentiment Classification
Dang Van Thin | Duong Ngoc Hao | Ngan Luu-Thuy Nguyen
Proceedings of the 38th Pacific Asia Conference on Language, Information and Computation
Dang Van Thin | Duong Ngoc Hao | Ngan Luu-Thuy Nguyen
Proceedings of the 38th Pacific Asia Conference on Language, Information and Computation
2023
ABCD Team at SemEval-2023 Task 12: An Ensemble Transformer-based System for African Sentiment Analysis
Dang Van Thin | Dai Ba Nguyen | Dang Ba Qui | Duong Ngoc Hao | Ngan Luu-Thuy Nguyen
Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)
Dang Van Thin | Dai Ba Nguyen | Dang Ba Qui | Duong Ngoc Hao | Ngan Luu-Thuy Nguyen
Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)
This paper describes the system of the ABCD team for three main tasks in the SemEval-2023 Task 12: AfriSenti-SemEval for Low-resource African Languages using Twitter Dataset. We focus on exploring the performance of ensemble architectures based on the soft voting technique and different pre-trained transformer-based language models. The experimental results show that our system has achieved competitive performance in some Tracks in Task A: Monolingual Sentiment Analysis, where we rank the Top 3, Top 2, and Top 4 for the Hause, Igbo and Moroccan languages. Besides, our model achieved competitive results and ranked $14ˆ{th}$ place in Task B (multilingual) setting and $14ˆ{th}$ and $8ˆ{th}$ place in Track 17 and Track 18 of Task C (zero-shot) setting.
2022
Search
Fix author
Co-authors
- Ngan Nguyen 4
- Nguyen Xuan Phuc 3
- Duong Ngoc Hao 2
- Nguyen Tuan Kiet 2
- Nguyen Xuan Phi 2
- Nguyen Van Tai 2
- Phan Minh Toan 2
- Nguyen Khanh Tuan Anh 1
- Cuong Xuan Chu 1
- Trong-Tai Dam Vu 1
- Christian Graf 1
- Le Minh Hung 1
- Tobias Kaminski 1
- Le Nguyen Anh Khoa 1
- An Nguyen Tran Khuong 1
- Khoa Anh-Nguyen Le 1
- Khoa Nguyen-Anh Le 1
- Tai Duc Le 1
- Huy Gia Luu 1
- Ngan Nguyen Luu-Thuy 1
- Nguyen Minh Bao 1
- Nguyen Minh Triet 1
- Ho Thuy Nga 1
- Hao Duong Ngoc 1
- Dai Ba Nguyen 1
- Dinh Khac Phuc Nguyen 1
- Duc-Vu Nguyen 1
- Long Hoang Huu Nguyen 1
- Nguyen Huu Dang Nguyen 1
- Thach Nguyen 1
- Vinh Nguyen 1
- Vinh Tiep Nguyen 1
- Xuan-Phi Nguyen 1
- Tram Nguyen Thi Ngoc 1
- Tran Nhan 1
- Nguyen Pham Hoang Le 1
- Duy Dang Phu 1
- Huynh Phu 1
- Tran Minh Quang 1
- Trang Minh Quang 1
- Dang Ba Qui 1
- Nguyen Minh Son 1
- Dam Vu Trong Tai 1
- Le Duc Tai 1
- Nguyễn Thiên Bảo 1
- Doan Nhat Tien 1
- Trung-Kien Tran 1
- Trinh Tran Tran 1
- Tram Truong 1
- Ho Thi Thanh Tuyen 1