Abdulkadir Bichi
Also published as: Abdulkadir Shehu Bichi, Abdulkadir Shehu Bichi
2026
VGU-M.Tech-AI at SemEval-2026: Multilingual Multi-Label Classification of Online Polarization Types via Weighted Transformer Fine-Tuning and Adaptive Per-Label Threshold Optimization
Abdulkadir Bichi | Jyoti Shekhawat
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Abdulkadir Bichi | Jyoti Shekhawat
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Abstract This research paper proposed a multilingual multi-label classification of online polarization types via weighted transformer fine-tuning and adaptive per-label threshold optimization (MMCOPT). Our task is to classify social media posts according to a given set of five labels. A post could be deemed to be politically, racially, religiously, or gender/sexually polarizing, or fall into the category of other. We incorporate a distilbert-base-multilingualcased model and attach a two-layer MLP head. We also use a class-imbalance-weighted binary cross-entropy loss and optimize thresholds for each class to improve the validation micro-F1 score. Our training set is drawn from the POLAR benchmark, the first large multilingual polarization dataset that includes posts from seven languages and multiple social media platforms. MMCOPT’s best internal validation micro-F1 score is 0.7855, and its macro-F1 score is 0.7749. Our model (team username: asbichi362) is ranked on the official Codabench leaderboard and shows competitive results across 22 language tracks of the research project multilingual polarization type classification, with its best results in Hindi (0.7429) and Urdu (0.7073).
Codezone Research Group at AraSentEval Shared Task: Arabic Sentiment Swap beyond Negation Prepending, Benchmarking Multilingual T5 against Large Language Models on the MA’AKS Corpus
Abdulkadir Shehu Bichi | Sarah Yassine
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Abdulkadir Shehu Bichi | Sarah Yassine
The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
Abstract We launched ASBN-MT5, the system for Arabic Sentiment Swap, which performs the task of inverting the sentiment of a sentence while keeping the meaning intact. This is a sequence-to-sequence task. We demonstrate ASBN-MT5: mT5, which is a MultiLingual T5 model, fine-tuned on the provided dataset of the AraSentEval 2026 Shared Task. We describe the data as the first of its kind for the Arabic language, as MAAKS is the first manually composed, parallel, cross-linguistic corpus for the Arabic language. With the preliminary results of Sentiment Flip for the task of Sentiment Inversion, we have recorded a rate of 59.5% for positive to negative conversions and 58.5% for negative to positive conversions, while maintaining an average similarity to the original sentences of 0.955. We present the Arabic prompts and a neuro-developmental (Deep Learning) recipe. Due to the evaluation criteria which include Exact Match, Flip Success, Surface Similarity, and Quality of Output, we restrict the use of Prepended Negation as the main technique and recommend the use of LLMs designed for the Arabic language in the near future. Keywords: mT5, sequence-to-sequence, AraSentEval 2026, Arabic NLP, Text Style Transfer, Sentiment Swap
HACS-TL: Cross-Script Transfer Learning for Hausa Ajami Hate Speech Detection Using Transformer-Based Architecture
Abdulkadir Shehu Bichi | Muqaddar Ali | Prashant Sharma | Ismail Dauda Abubakar
Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script
Abdulkadir Shehu Bichi | Muqaddar Ali | Prashant Sharma | Ismail Dauda Abubakar
Proceedings of the 2nd Workshop on NLP for Languages Using Arabic Script
The Arabic-derived scripts contain several languages that face challenges with the limited resources of speech detection, these challenges are worsened by the scarcity of resources and highly complex linguistic challenges. We proposed ( HACS-TL Hausa Ajami Cross-Script Transfer Learning) a brand new transformer-based architecture that focuses on the detection of hate speech within Ajami script. Hausa is a Chadic language which contains over 77 million speakers located in West Africa; it uses two types of scripts: the Latin (Boko) and the Arabic-derived Ajami which creates new computational difficulties. Our method combines scripts of artistically converted linguistics, augmented cross script multi-head attention, and dialect feature extraction to trellis the morphophonological depth of the Hausa. After a thorough examination using stratified cross-validation along with systemically augmented data, HACS-TL obtained a Macro F1 score of 76.09% which is a significant improvement from the other multilingual baselines (mBERT (69.17 % ) XLM-RoBERTa (73.20 % ) AraBERT (58.63% ) ) HACS-TL outperformed all of the previously stated models. Strong multilingual baselines refer to the other stated models; AraBERT (58.63) XLM-RoBERTa (73.20) mBERT (69.17) HACS-TL 70.73 + 10 % Cross-Script+ (mBERT) 46.73 + 0.9 % Cross-Script + AraBERT. The importance of cross-script attention and learning from transfer sources of resources to languages with limited scripts has proven effective. Our systematic method has aided the advancement of Arabic script homage Hausa and African language resources for the NLP of the Nubians in learning African languages and the intricate Nubian and cross-learning systems from different scripts.
2025
HausaNLP at SemEval-2025 Task 3: Towards a Fine-Grained Model-Aware Hallucination Detection
Maryam Bala | Amina Abubakar | Abdulhamid Abubakar | Abdulkadir Bichi | Hafsa Ahmad | Sani Abdullahi Sani | Idris Abdulmumin | Shamsuddeen Hassan Muhammad | Ibrahim Said Ahmad
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Maryam Bala | Amina Abubakar | Abdulhamid Abubakar | Abdulkadir Bichi | Hafsa Ahmad | Sani Abdullahi Sani | Idris Abdulmumin | Shamsuddeen Hassan Muhammad | Ibrahim Said Ahmad
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
This paper presents our findings of the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes, MU-SHROOM, which focuses on identifying hallucinations and related overgeneration errors in large language models (LLMs). The shared task involves detecting specific text spans that constitute hallucinations in the outputs generated by LLMs in 14 languages. To address this task, we aim to provide a nuanced, model-aware understanding of hallucination occurrences and severity in English. We used natural language inference and fine-tuned a ModernBERT model using a synthetic dataset of 400 samples, achieving an Intersection over Union (IoU) score of 0.032 and a correlation score of 0.422. These results indicate a moderately positive correlation between the model’s confidence scores and the actual presence of hallucinations. The IoU score indicates that our modelhas a relatively low overlap between the predicted hallucination span and the truth annotation. The performance is unsurprising, given the intricate nature of hallucination detection. Hallucinations often manifest subtly, relying on context, making pinpointing their exact boundaries formidable.