Yasir Zaki
2026
Analyzing Political Stances on Twitter/X in the Lead-up to the 2024 U.S. Election
Hazem Ibrahim | Farhan Kamrul Khan | Yasir Zaki | Talal Rahwan
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Hazem Ibrahim | Farhan Kamrul Khan | Yasir Zaki | Talal Rahwan
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Social media platforms play a pivotal role in shaping public opinion and amplifying political discourse, particularly during elections. However, the same dynamics that foster democratic engagement can also exacerbate polarization. To better understand these challenges, here, we investigate the ideological positioning of tweets related to the 2024 U.S. Presidential Election. To this end, we analyze 1,235 tweets from key political figures and 63,322 replies, and classify ideological stances into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. Using a classification pipeline involving three large language models (LLMs)—GPT-4o, Gemini-Pro, and Claude-Opus—and validated by human annotators, we explore how ideological alignment varies between candidates and constituents. We find that Republican candidates author significantly more tweets in criticism of the Democratic party and its candidates than vice versa, but this relationship does not hold for replies to candidate tweets. Furthermore, we highlight shifts in public discourse observed during key political events. By shedding light on the ideological dynamics of online political interactions, these results provide insights for policymakers and platforms seeking to address polarization and foster healthier political dialogue.
2025
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
Essa Jan | Moiz Ali | Muhammad Saram Hassan | Muhammad Fareed Zaffar | Yasir Zaki
Findings of the Association for Computational Linguistics: EMNLP 2025
Essa Jan | Moiz Ali | Muhammad Saram Hassan | Muhammad Fareed Zaffar | Yasir Zaki
Findings of the Association for Computational Linguistics: EMNLP 2025
As the knowledge of large language models (LLMs) becomes outdated over time, there is a growing need for efficient methods to update them, especially when injecting proprietary information. Our study reveals that comprehension-intensive fine-tuning tasks (e.g., question answering and blanks) achieve substantially higher knowledge retention rates (48%) compared to mapping-oriented tasks like translation (17%) or text-to-JSON conversion (20%), despite exposure to identical factual content. We demonstrate that this pattern persists across model architectures and follows scaling laws, with larger models showing improved retention across all task types. However, all models exhibit significant performance drops when applying injected knowledge in broader contexts, suggesting limited semantic integration. These findings show the importance of task selection in updating LLM knowledge, showing that effective knowledge injection relies not just on data exposure but on the depth of cognitive engagement during fine-tuning.
Multitask-Bench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
Essa Jan | Nouar Aldahoul | Moiz Ali | Faizan Ahmad | Fareed Zaffar | Yasir Zaki
Proceedings of the 31st International Conference on Computational Linguistics
Essa Jan | Nouar Aldahoul | Moiz Ali | Faizan Ahmad | Fareed Zaffar | Yasir Zaki
Proceedings of the 31st International Conference on Computational Linguistics
Recent breakthroughs in Large Language Models (LLMs) have led to their adoption across a wide range of tasks, ranging from code generation to machine translation and sentiment analysis, etc. Red teaming/Safety alignment efforts show that fine-tuning models on benign (non-harmful) data could compromise safety. However, it remains unclear to what extent this phenomenon is influenced by different variables, including fine-tuning task, model calibrations, etc. This paper explores the task-wise safety degradation due to fine-tuning on downstream tasks such as summarization, code generation, translation, and classification across various calibration. Our results reveal that: 1) Fine-tuning LLMs for code generation and translation leads to the highest degradation in safety guardrails. 2) LLMs generally have weaker guardrails for translation and classification, with 73-92% of harmful prompts answered, across baseline and other calibrations, falling into one of two concern categories. 3) Current solutions, including guards and safety tuning datasets, lack cross-task robustness. To address these issues, we developed a new multitask safety dataset effectively reducing attack success rates across a range of tasks without compromising the model’s overall helpfulness. Our work underscores the need for generalized alignment measures to ensure safer and more robust models.
NYUAD at MAHED Shared Task: Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
Nouar AlDahoul | Yasir Zaki
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Nouar AlDahoul | Yasir Zaki
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks