Gabor Proszeky
Also published as: Gábor Prószéky, Gabor Prbszeky
2026
Managing Growth in a National Corpus: The Hungarian National Corpus 3.0 (MNSZ3)
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
The third generation of the Hungarian National Corpus (MNSZ3) aims to provide a large-scale, curated, and well-described corpus resource needed for the sustainable digital presence of Hungarian. Building on the domain structure and proportions of MNSZ2 (v2.0.5; 1.04 billion running words), the project targets a substantial increase in scale while also strengthening the coverage and metadata description of Hungarian language use outside Hungary. MNSZ3 retains the six traditional domains of the earlier corpus—press, fiction, scientific, official, personal, and transcribed spoken language—and is planned to reach approximately 10 billion tokens. This paper presents the motivation and design principles of the project, outlines the practical decisions and procedures used in data collection and cleaning, and discusses the annotation strategy developed for large-scale processing. In planning the linguistic analysis, we build on the complementary strengths of HuSpaCy and e-magyar: HuSpaCy provides the unified and efficient UD-oriented processing backbone, while e-magyar (emMorph) is preserved as an explicit additional layer for morphology and lemmatisation.
2025
HuGME: A benchmark system for evaluating Hungarian generative LLMs
Noémi Ligeti-Nagy | Gabor Madarasz | Flora Foldesi | Mariann Lengyel | Matyas Osvath | Bence Sarossy | Kristof Varga | Győző Zijian Yang | Enikő Héja | Tamás Váradi | Gábor Prószéky
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
Noémi Ligeti-Nagy | Gabor Madarasz | Flora Foldesi | Mariann Lengyel | Matyas Osvath | Bence Sarossy | Kristof Varga | Győző Zijian Yang | Enikő Héja | Tamás Váradi | Gábor Prószéky
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
In this study, we introduce the Hungarian Generative Model Evaluation (HuGME) benchmark, a new framework designed to assess the linguistic proficiency of large language models (LLMs) in Hungarian. HuGME evaluates models across a diverse set of linguistic and reasoning skills, including bias, toxicity, faithfulness, relevance, summarization, prompt alignment, readability, spelling, grammaticality, and domain-specific knowledge through tasks like TruthfulQA and MMLU. We applied HuGME to a range of Hungarian LLMs, including those developed in-house as well as several publicly available models that claim Hungarian language proficiency. This paper presents the comparative results of these evaluations, shedding light on the capabilities of current LLMs in processing the Hungarian language. Through our analysis, we aim to both showcase the current state of Hungarian linguistic processing in LLMs and provide a foundational resource for future advancements in the field.
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
Haote Yang | Xingjian Wei | Jiang Wu | Noémi Ligeti-Nagy | Jiaxing Sun | Yinfan Wang | Zijian Győző Yang | Junyuan Gao | Jingchao Wang | Bowen Jiang | Shasha Wang | Nanjun Yu | Zihao Zhang | Shixin Hong | Hongwei Liu | Wei Li | Songyang Zhang | Dahua Lin | Lijun Wu | Gábor Prószéky | Conghui He
Findings of the Association for Computational Linguistics: ACL 2025
Haote Yang | Xingjian Wei | Jiang Wu | Noémi Ligeti-Nagy | Jiaxing Sun | Yinfan Wang | Zijian Győző Yang | Junyuan Gao | Jingchao Wang | Bowen Jiang | Shasha Wang | Nanjun Yu | Zihao Zhang | Shixin Hong | Hongwei Liu | Wei Li | Songyang Zhang | Dahua Lin | Lijun Wu | Gábor Prószéky | Conghui He
Findings of the Association for Computational Linguistics: ACL 2025
We introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. OpenHuEval is constructed from a vast collection of Hungarian-specific materials sourced from multiple origins. In the construction, we incorporated the latest design principles for evaluating LLMs, such as using real user queries from the internet, emphasizing the assessment of LLMs’ generative capabilities, and employing LLM-as-judge to enhance the multidimensionality and accuracy of evaluations. Ultimately, OpenHuEval encompasses eight Hungarian-specific dimensions, featuring five tasks and 3953 questions. Consequently, OpenHuEval provides the comprehensive, in-depth, and scientifically accurate assessment of LLM performance in the context of the Hungarian language and its specifics. We evaluated current mainstream LLMs, including both traditional LLMs and recently developed Large Reasoning Models. The results demonstrate the significant necessity for evaluation and model optimization tailored to the Hungarian language and specifics. We also established the framework for analyzing the thinking processes of LRMs with OpenHuEval, revealing intrinsic patterns and mechanisms of these models in non-English languages, with Hungarian serving as a representative example. We will release OpenHuEval at https://github.com/opendatalab/OpenHuEval .
2014
Almost fifty years after the (first?) ALPAC report
Gábor Prószéky
Proceedings of Translating and the Computer 36
Gábor Prószéky
Proceedings of Translating and the Computer 36
2011
Endangered Uralic Languages and Language Technologies
Gábor Prószéky
Proceedings of the Workshop on Language Technologies for Digital Humanities and Cultural Heritage
Gábor Prószéky
Proceedings of the Workshop on Language Technologies for Digital Humanities and Cultural Heritage
2008
The MetaMorpho Translation System
Attila Novák | László Tihanyi | Gábor Prószéky
Proceedings of the Third Workshop on Statistical Machine Translation
Attila Novák | László Tihanyi | Gábor Prószéky
Proceedings of the Third Workshop on Statistical Machine Translation
2005
An approach to machine translation via the rule-to-rule hypothesis
Gábor Prószéky
Proceedings of the 10th EAMT Conference: Practical applications of machine translation
Gábor Prószéky
Proceedings of the 10th EAMT Conference: Practical applications of machine translation
2004
Moose: a robust high-performance parser and generator
Gábor Prószéky | László Tihanyi | Gábor Ugray
Proceedings of the 9th EAMT Workshop: Broadening horizons of machine translation and its applications
Gábor Prószéky | László Tihanyi | Gábor Ugray
Proceedings of the 9th EAMT Workshop: Broadening horizons of machine translation and its applications
2003
Annotated Hungarian National Corpus
Zoltán Alexin | János Csirik | Tibor Gyimóthy | Károly Bibok | Csaba Hatvani | Gábor Prószéky | László Tihanyi
10th Conference of the European Chapter of the Association for Computational Linguistics
Zoltán Alexin | János Csirik | Tibor Gyimóthy | Károly Bibok | Csaba Hatvani | Gábor Prószéky | László Tihanyi
10th Conference of the European Chapter of the Association for Computational Linguistics
2002
Automatism and User Interaction: Building a Hungarian WordNet
Gábor Prószéky | Márton Miháltz
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Gábor Prószéky | Márton Miháltz
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Context-Sensitive Electronic Dictionaries
Gábor Prószéky | Balázs Kis
COLING 2002: The 17th International Conference on Computational Linguistics: Project Notes
Gábor Prószéky | Balázs Kis
COLING 2002: The 17th International Conference on Computational Linguistics: Project Notes
Recognition Assistance - Treating Errors in Texts Acquired from Various Recognition Processes
Gábor Prószéky | Mátyás Naszódi | Balázs Kis
COLING 2002: The 17th International Conference on Computational Linguistics: Project Notes
Gábor Prószéky | Mátyás Naszódi | Balázs Kis
COLING 2002: The 17th International Conference on Computational Linguistics: Project Notes
MetaMorpho: A Pattern-Based Machine Translation System
Gábor Prószéky
Proceedings of Translating and the Computer 24
Gábor Prószéky
Proceedings of Translating and the Computer 24
1999
A Unification-based Approach to Morpho-syntactic Parsing of Agglutinative and Other (Highly) Inflectional Languages
Gabor Proszeky | Balazs Kis
Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics
Gabor Proszeky | Balazs Kis
Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics
1998
An Intelligent Multi-Dictionary Environment
Gabor Prbszeky
36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, Volume 2
Gabor Prbszeky
36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, Volume 2
An Intelligent Multi-Dictionary Environment
Gabor Proszeky
COLING 1998 Volume 2: The 17th International Conference on Computational Linguistics
Gabor Proszeky
COLING 1998 Volume 2: The 17th International Conference on Computational Linguistics
1997
Reading more into Foreign Languages
John Nerbonne | Lauri Karttunen | Elena Paskaleva | Gabor Proszeky | Tiit Roosmaa
Fifth Conference on Applied Natural Language Processing
John Nerbonne | Lauri Karttunen | Elena Paskaleva | Gabor Proszeky | Tiit Roosmaa
Fifth Conference on Applied Natural Language Processing
1996
Morphological Analyzer as Syntactic Parser
Gábor Prószéky
COLING 1996 Volume 2: The 16th International Conference on Computational Linguistics
Gábor Prószéky
COLING 1996 Volume 2: The 16th International Conference on Computational Linguistics
1994
Humor-Based Applications
Gabor Proszeky | Miklos Pal | Laszlo Tihanyi
COLING 1994 Volume 2: The 15th International Conference on Computational Linguistics
Gabor Proszeky | Miklos Pal | Laszlo Tihanyi
COLING 1994 Volume 2: The 15th International Conference on Computational Linguistics
Industrial Applications of Unification Morphology
Gabor Proszeky
Fourth Conference on Applied Natural Language Processing
Gabor Proszeky
Fourth Conference on Applied Natural Language Processing
1993
Helyette: Inflectional Thesaurus for Agglutinative Languages
Gabor Proszeky | Laszlo Tihanyi
Sixth Conference of the European Chapter of the Association for Computational Linguistics
Gabor Proszeky | Laszlo Tihanyi
Sixth Conference of the European Chapter of the Association for Computational Linguistics
1986
Search
Fix author
Co-authors
- Laszlo Tihanyi 5
- Balázs Kis 3
- Noémi Ligeti-Nagy 3
- Flóra Földesi 2
- Enikő Héja 2
- Bence Sárossy 2
- Tamás Váradi 2
- Zoltán Alexin 1
- Károly Bibok 1
- Ágnes Bánfi 1
- János Csirik 1
- Junyuan Gao 1
- Tibor Gyimóthy 1
- Csaba Hatvani 1
- Conghui He 1
- Shixin Hong 1
- Bowen Jiang 1
- Lauri Karttunen 1
- Mariann Lengyel 1
- Wei Li 1
- Dahua Lin 1
- Hongwei Liu 1
- Gabor Madarasz 1
- Márton Miháltz 1
- Mátyás Naszódi 1
- John Nerbonne 1
- Attila Novák 1
- Mátyás Osváth 1
- Miklos Pal 1
- Elena Paskaleva 1
- Tiit Roosmaa 1
- Boglárka Skrabák 1
- Jiaxing Sun 1
- Gábor Ugray 1
- Kristof Varga 1
- Jingchao Wang 1
- Shasha Wang 1
- Yinfan Wang 1
- Xingjian Wei 1
- Jiang Wu 1
- Lijun Wu 1
- Győző Zijian Yang 1
- Haote Yang 1
- Zijian Győző Yang 1
- Nanjun Yu 1
- Songyang Zhang 1
- Zihao Zhang 1