Sašo Džeroski
2026
FoodBench-QA: Overview of the Shared Task on Grounded Food and Nutrition Question Answering
Tome Eftimov | Ana Gjorgjevikj | Matej Martinc | Gjorgjina Cenikj | Sašo Džeroski | Barbara Koroušič Seljak
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
Tome Eftimov | Ana Gjorgjevikj | Matej Martinc | Gjorgjina Cenikj | Sašo Džeroski | Barbara Koroušič Seljak
Proceedings of the Third Workshop on Patient-Oriented Language Processing (CL4Health) @ LREC 2026
We present the results of the FoodBench-QA 2026 shared task at the CL4Health workshop, collocated with LREC 2026. FoodBench-QA challenges systems to answer food and nutrition questions using evidence from food composition databases and food-related ontologies. The shared task comprises three main tasks: nutrient estimation from recipe ingredients, evaluated using EU Regulation 1169/2011 tolerance thresholds; FSA traffic-light classification for fat, salt, saturates, and sugars; and food named entity recognition and linking to three ontologies, namely Hansard Taxonomy, FoodOn, and SNOMED CT. We received submissions from five participating teams across all tasks. For nutrient estimation, the best system achieved accuracy rates of 93.57% for protein, 86.50% for sugars, 84.65% for fat, and 86.26% for saturates. For FSA traffic-light prediction, the best macro F1 scores ranged from 0.65 to 0.90 across different nutrient-color combinations. For named entity linking, the best systems achieved macro F1 scores between 60.71% and 80.89% for natural text and 87.75% and 95.75% for artificial NEL datasets, depending on the ontology.
2006
Towards a Slovene Dependency Treebank
Sašo Džeroski | Tomaž Erjavec | Nina Ledinek | Petr Pajas | Zdenek Žabokrtsky | Andreja Žele
Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06)
Sašo Džeroski | Tomaž Erjavec | Nina Ledinek | Petr Pajas | Zdenek Žabokrtsky | Andreja Žele
Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06)
The paper presents the initial release of the Slovene Dependency Treebank, currently containing 2000 sentences or 30.000 words. Our approach to annotation is based on the Prague Dependency Treebank, which serves as an excellent model due to the similarity of the languages, the existence of a detailed annotation guide and an annotation editor. The initial treebank contains a portion of the MULTEXT-East parallel word-level annotated corpus, namely the first part of the Slovene translation of Orwell’s “1984”. This corpus was first parsed automatically, to arrive at the initial analytic level dependency trees. These were then hand corrected using the tree editor TrEd; simultaneously, the Czech annotation manual was modified for Slovene. The current version is available in XML/TEI, as well as derived formats, and has been used in a comparative evaluation using the MALT parser, and as one of the languages present in the CoNLL-X shared task on dependency parsing. The paper also discusses further work, in the first instance the composition of the corpus to be annotated next.
2002
A Machine Learning Approach to Automatic Functor Assignment in the Prague Dependency Treebank
Zdeněk Žabokrtský | Petr Sgall | Sašo Džeroski
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Zdeněk Žabokrtský | Petr Sgall | Sašo Džeroski
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)