A Multilingual Linguistic Analysis of Human vs LLM-Generated News in a Disinformation Context
Silvia Gargova, Alba Perez-Montero, Elena Lloret Pastor, Paloma Moreda Pozo
Abstract
The rise of Large Language Models has shifted the Information Disorder landscape toward automated threats. This study investigates the linguistic construction of synthetic news by comparing GPT-5, Gemini 2.5, and Grok 4 across English, Spanish, and Bulgarian. Using multilingual human-authored verified news and disinformation as seeds, we analyze how prompt informativeness and model architecture influence deceptive content production. Our methodology employs five metrics: semantic similarity, factual consistency, readability, lexical richness, and persuasion technique frequency. Our analysis reveals that while prompt scarcity leads to informational loss, LLMs maintain a homogenized stylistic template regardless of input length. Unlike human authors, who intensify rhetorical and emotional markers to drive deceptive intent, LLMs adhere to a neutral register. This study identifies distinct statistical patterns in generated content characterized by hyper-standardized readability and high lexical density (p < 0.001). These features serve as robust “LLM signatures”, enabling a classification accuracy of 96% across English, Spanish, and Bulgarian. These findings suggest that generated disinformation relies on invariant syntactic structures rather than nuanced human rhetoric, providing a framework for detection tools centered on structural patterns rather than content veracity.- Anthology ID:
- 2026.indor-1.9
- Volume:
- Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Editors:
- Simona Frenda, Marco Antonio Stranisci, Shaina Ashraf, Ada Ren, Ioannis Konstas, Usman Naseem
- Venues:
- InDor | WS
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- Note:
- Pages:
- 85–96
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-indor-09
- DOI:
- 10.63317/345wj5pp9nkh
- Cite (ACL):
- Silvia Gargova, Alba Perez-Montero, Elena Lloret Pastor, and Paloma Moreda Pozo. 2026. A Multilingual Linguistic Analysis of Human vs LLM-Generated News in a Disinformation Context. In Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026, pages 85–96, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
- Cite (Informal):
- A Multilingual Linguistic Analysis of Human vs LLM-Generated News in a Disinformation Context (Gargova et al., InDor 2026)