Maryam Mohammadi
2026
Do Large Language Models Understand Double Mismatches? Evidence from Farsi
Maryam Mohammadi
The Proceedings of the First Workshop on NLP and LLMs for the Iranian Language Family
Maryam Mohammadi
The Proceedings of the First Workshop on NLP and LLMs for the Iranian Language Family
Large language models (LLMs) are increasingly used for communication in many languages, therefore, understanding their limitations with respect to culture-specific pragmatics is important. While LLMs perform well on statistically frequent structures, their shortcomings are most evident in rare pragmatic phenomena. This study investigates whether LLMs can generate a (rare) complex honorific mismatch in Farsi. The pattern arises at two levels:(i) a plural pronoun disagrees with a singular referent for the sake of honorification, and (ii) the related components violate the Polite Plural Generalization due to intimacy implication. This double mismatch pattern is attested in everyday speech, though it is statistically sparse. We tested GPT-4 across multiple scenarios. The results reveal that the model successfully employs the first mismatch to indicate honorific, but fails to adopt the second mismatch that simultaneously conveys intimacy. The model thus deviates from humanlike behavior at the syntax–pragmatics interface. These findings suggest that, while machine models demonstrate partial success in generating honorifics, they rely primarily on statistical patterns and lack the deeper pragmatic understanding necessary for contextual competence.
Modelling Legal Compliance in a Consent Wizard Application as Part of a Research-Centered and User-Oriented Data Infrastructure
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Recent research calls for data management infrastructures that explicitly operate within the bounds of ethical and legal constraints, and facilitate adherence to Open Science principles by integrating automated support for planning, collection, storage, use, reuse, and sharing of data within. Legal and ethical requirements of data processing have become increasingly complex, introducing administrative barriers to scientific research investigating data generated by human participants, which encompasses a vast majority of humanities research. In response to this, we present RUDI (“Research-centered User-oriented Data Infrastructure”), a modular framework grounded in an interdisciplinary approach informed by legal, computational and linguistic expertise. This paper introduces its first component; a configurable and dynamically adaptive consent form generator in the form of a “wizard” web application. We outline how legal aspects are modeled within, and highlight its concrete benefits for administrative aspects of research. Further, we discuss the contextualization of data within the research domain by leveraging the use of standardized ontology within the framework.
2023
mage as a bias particle in interrogatives
Maryam Mohammadi
Proceedings of the 4th Workshop on Inquisitiveness Below and Beyond the Sentence Boundary
Maryam Mohammadi
Proceedings of the 4th Workshop on Inquisitiveness Below and Beyond the Sentence Boundary
This paper investigates Farsi particle ‘mage’ in interrogatives, including both polar and constituent/Wh questions. I will show that ‘mage’ requires both contextual evidence and speaker’s prior belief in the sense that they contradict each other. While in polar questions (PQs) both types of bias can be straightforwardly expressed through the uttered proposition (cf. Mameni 2010), Wh-questions (WhQs) do not provide such a propositional object. To capture this difference, I propose Answerhood as the relevant notation that provides the necessary object source for ‘mage’ (inspired by Theiler 2021). The proposal establishes the felicity conditions and the meaning of ‘mage’ in relation to the (contextually) restricted answerhood in both polar and constituent questions.