Sachi Kato
2026
Automatic Detection of Metaphorical Expressions in Classical Japanese Using WLSP-Enhanced BERT
Hang Zhu | Mitoki Ohara | Rei Kikuchi | Kanako Komiya | Masayuki Asahara | Sachi Kato
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Hang Zhu | Mitoki Ohara | Rei Kikuchi | Kanako Komiya | Masayuki Asahara | Sachi Kato
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Metaphor detection is a fundamental task in natural language processing, yet research on historical languages remains limited. While progress has been made in modern Japanese metaphor detection, classical Japanese texts present unique challenges due to their distinct vocabulary, grammar, and metaphorical patterns. This paper addresses this gap by applying a BERT-based metaphor detection method enhanced with semantic classification information from the Word List by Semantic Principles (WLSP) to classical Japanese texts. We evaluate our approach on CHJ-Metaphor, a newly available corpus featuring metaphor annotations for three medieval Japanese works from the Corpus of Historical Japanese (CHJ). Our method achieves an F1-score of 82.18 through 5-fold cross-validation. Notably, qualitative analysis by domain experts reveals that our model successfully identifies genuine metaphors overlooked during manual annotation, demonstrating its potential as a tool for improving annotation quality in large-scale corpus construction. These results confirm the effectiveness of WLSP-enhanced approaches for metaphor detection in classical Japanese and suggest promising directions for applying similar techniques to other historical languages.
2025
Large-Scale Japanese Metaphor Corpus Construction: Expanding BCCWJ-Metaphor with Automated Annotation
Hang Zhu | Rowan Hall Maudslay | Kanako Komiya | Sachi Kato | Masayuki Asahara
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
Hang Zhu | Rowan Hall Maudslay | Kanako Komiya | Sachi Kato | Masayuki Asahara
Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation
2024
Assigning Impression Rating Information to the ‘Balanced Corpus of Contemporary Written Japanese’
Sachi Kato | Masayuki Asahara
Proceedings of the 38th Pacific Asia Conference on Language, Information and Computation
Sachi Kato | Masayuki Asahara
Proceedings of the 38th Pacific Asia Conference on Language, Information and Computation
2022
CHJ-WLSP: Annotation of ‘Word List by Semantic Principles’ Labels for the Corpus of Historical Japanese
Masayuki Asahara | Nao Ikegami | Tai Suzuki | Taro Ichimura | Asuko Kondo | Sachi Kato | Makoto Yamazaki
Proceedings of the Second Workshop on Language Technologies for Historical and Ancient Languages
Masayuki Asahara | Nao Ikegami | Tai Suzuki | Taro Ichimura | Asuko Kondo | Sachi Kato | Makoto Yamazaki
Proceedings of the Second Workshop on Language Technologies for Historical and Ancient Languages
This article presents a word-sense annotation for the Corpus of Historical Japanese: a mashed-up Japanese lexicon based on the ‘Word List by Semantic Principles’ (WLSP). The WLSP is a large-scale Japanese thesaurus that includes 98,241 entries with syntactic and hierarchical semantic categories. The historical WLSP is also compiled for the words in ancient Japanese. We utilized a morpheme-word sense alignment table to extract all possible word sense candidates for each word appearing in the target corpus. Then, we manually disambiguated the word senses for 647,751 words in the texts from the 10th century to 1910.
2021
The Annotation of Antonym Information in the ‘Word List by Semantic Principles’
Sachi Kato | Masayuki Asahara | Nanami Moriyama | Makoto Yamazaki Asami Ogiwara
Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation
Sachi Kato | Masayuki Asahara | Nanami Moriyama | Makoto Yamazaki Asami Ogiwara
Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation
2018
Annotation of ‘Word List by Semantic Principles’ Labels for the Balanced Corpus of Contemporary Written Japanese
Sachi Kato | Masayuki Asahara | Makoto Yamazaki
Proceedings of the 32nd Pacific Asia Conference on Language, Information and Computation
Sachi Kato | Masayuki Asahara | Makoto Yamazaki
Proceedings of the 32nd Pacific Asia Conference on Language, Information and Computation
2017
Between Reading Time and Syntactic/Semantic Categories
Masayuki Asahara | Sachi Kato
Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
Masayuki Asahara | Sachi Kato
Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
This article presents a contrastive analysis between reading time and syntactic/semantic categories in Japanese. We overlaid the reading time annotation of BCCWJ-EyeTrack and a syntactic/semantic category information annotation on the ‘Balanced Corpus of Contemporary Written Japanese’. Statistical analysis based on a mixed linear model showed that verbal phrases tend to have shorter reading times than adjectives, adverbial phrases, or nominal phrases. The results suggest that the preceding phrases associated with the presenting phrases promote the reading process to shorten the gazing time.
2016
‘BonTen’ – Corpus Concordance System for ‘NINJAL Web Japanese Corpus’
Masayuki Asahara | Kazuya Kawahara | Yuya Takei | Hideto Masuoka | Yasuko Ohba | Yuki Torii | Toru Morii | Yuki Tanaka | Kikuo Maekawa | Sachi Kato | Hikari Konishi
Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: System Demonstrations
Masayuki Asahara | Kazuya Kawahara | Yuya Takei | Hideto Masuoka | Yasuko Ohba | Yuki Torii | Toru Morii | Yuki Tanaka | Kikuo Maekawa | Sachi Kato | Hikari Konishi
Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: System Demonstrations
The National Institute for Japanese Language and Linguistics, Japan (NINJAL) has undertaken a corpus compilation project to construct a web corpus for linguistic research comprising ten billion words. The project is divided into four parts: page collection, linguistic analysis, development of the corpus concordance system, and preservation. This article presents the corpus concordance system named ‘BonTen’ which enables the ten-billion-scaled corpus to be queried by string, a sequence of morphological information or a subtree of the syntactic dependency structure.