Doria Bonzi
2026
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training
Doria Bonzi | Tom Bourgeade | Fabrice Lefèvre | Irina Illina
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Doria Bonzi | Tom Bourgeade | Fabrice Lefèvre | Irina Illina
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consist of brief scenario-driven simulations of doctor-patient interactions. However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs). To address this gap, we introduce a French OSCE dialogue dataset comprising 240 real student–patient training interactions. We build upon it a controllable LLM-based pipeline to generate synthetic OSCE dialogues. The pipeline integrates modular components, such as retrieval-based grounding and a reflection loop, to ensure patient fidelity, coherence, and realism. Additionally, we propose a multi-level evaluation framework assessing patient simulation quality, student performance, and linguistic quality, using an LLM-as-a-Judge approach. Experiments suggest that controllability modules generally improve patient fidelity and student evaluation consistency. Finally, we also implement an interactive prototype in which students can practice with a VP and receive automatic feedback.
CareMedEval Dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
Doria Bonzi | Alexandre Guiggi | Frederic Bechet | Carlos Ramisch | Benoit Favre
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Doria Bonzi | Alexandre Guiggi | Frederic Bechet | Carlos Ramisch | Benoit Favre
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for critical reasoning in specialized domains. We introduce CareMedEval, an original dataset designed to evaluate LLMs on biomedical critical appraisal and reasoning tasks. Derived from authentic exams taken by French medical students, the dataset contains 534 questions based on 37 scientific articles. Unlike existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in scientific papers. Benchmarking state-of-the-art generalist and biomedical-specialized LLMs under various context conditions reveals the difficulty of the task: open and commercial models fail to exceed an Exact Match Rate of 0.5 even though generating intermediate reasoning tokens considerably improves the results. Yet, models remain challenged especially on questions about study limitations and statistical analysis. CareMedEval provides a challenging benchmark for grounded reasoning, exposing current LLM limitations and paving the way for future development of automated support for critical appraisal.