When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis
Mohamed Ibrahim, Abdullah Makki, Youssef Barakat, Nour Samy, Sarah AlHumoud
Abstract
This study evaluates the performance of a fine-tuned Arabic sentiment transformer (CAMeL-MSA) against eight large language models (LLMs). Using zero-shot prompting across six Arabic sentiment datasets, we compare a specialized, task-specific approach against generalized model capabilities. Results show that the fine-tuned baseline substantially outperformed all LLMs on five of the six datasets in both accuracy and Macro F1-score. While LLMs offer versatility, this comparison highlights the continued practical superiority of task-specific fine-tuning over zero-shot prompting.- Anthology ID:
- 2026.osact-1.4
- Volume:
- The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
- Month:
- May
- Year:
- 2026
- Address:
- Palma, Mallorca (Spain)
- Editors:
- Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
- Venues:
- OSACT | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 35–39
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-osact-04
- DOI:
- 10.63317/2kimof4u6y8x
- Cite (ACL):
- Mohamed Ibrahim, Abdullah Makki, Youssef Barakat, Nour Samy, and Sarah AlHumoud. 2026. When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 35–39, Palma, Mallorca (Spain). Association for Computational Linguistics.
- Cite (Informal):
- When Bigger Isn’t Better: Evaluating LLMs for Arabic Sentiment Analysis (Ibrahim et al., OSACT 2026)