Seunghyeok Hong
2026
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
Hanwool Lee | Dasol Choi | Sooyong Kim | Ilgyun Jung | Sangwon Baek | Guijin Son | Inseong Hwang | Naeun Lee | Seunghyeok Hong
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Hanwool Lee | Dasol Choi | Sooyong Kim | Ilgyun Jung | Sangwon Baek | Guijin Son | Inseong Hwang | Naeun Lee | Seunghyeok Hong
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps does not mean enforcing a one-size-fits-all evaluation. Rather, effective benchmarking requires diverse experimental approaches and a framework robust enough to support them. To this end, we introduce HRET (Haerae Evaluation Toolkit), an open-source, registry-based framework that unifies Korean LLM assessment. HRET integrates major Korean benchmarks, multiple inference backends, and multi-method evaluation, with language consistency enforcement to ensure genuine Korean outputs. Its modular registry design also enables rapid incorporation of new datasets, methods, and backends, ensuring the toolkit adapts to evolving research needs. Beyond standard accuracy metrics, HRET incorporates Korean-focused output analyses-morphology-aware Type-Token Ratio (TTR) for evaluating lexical diversity and systematic keyword-omission detection for identifying missing concepts-to provide diagnostic insights into language-specific behaviors. These targeted analyses help researchers pinpoint morphological and semantic shortcomings in model outputs, guiding focused improvements in Korean LLM development.
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
Dasol Choi | Guijin Son | Hanwool Lee | Minhyuk Kim | Hyunwoo Ko | Teabin Lim | Eungyeol Ahn | Jungwhan Kim | Seunghyeok Hong | Youngsook Song
Findings of the Association for Computational Linguistics: ACL 2026
Dasol Choi | Guijin Son | Hanwool Lee | Minhyuk Kim | Hyunwoo Ko | Teabin Lim | Eungyeol Ahn | Jungwhan Kim | Seunghyeok Hong | Youngsook Song
Findings of the Association for Computational Linguistics: ACL 2026
Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. Users naturally leave much unsaid, relying on images to convey context. We introduce HAERAE-Vision, a benchmark of 653 real-world visual questions from Korean online communities (0.76% survival from 86K candidates), each paired with an explicit rewrite, yielding 1,306 query variants in total. Evaluating 39 VLMs, we find that even state-of-the-art models (GPT-5, Gemini 2.5 Pro) achieve under 50% on the original queries. Crucially, query explicitation alone yields 8 to 22 point improvements, with smaller models benefiting most. We further show that even with web search, under-specified queries underperform explicit queries without search, revealing that current retrieval cannot compensate for what users leave unsaid. Our findings demonstrate that a substantial portion of VLM difficulty stem from natural query under-specification instead of model capability, highlighting a critical gap between benchmark evaluation and real-world deployment.