Abstract
We propose methods for transliterating English loanwords in Japanese from their Japanese written form (katakana/romaji) to their original English written form. Our data is a Japanese-English loanwords dictionary that we have created ourselves. We employ two approaches: direct transliteration, which directly converts words from katakana to English, and indirect transliteration, which utilizes the English pronunciation as a means to convert katakana words into their corresponding English sound representations, which are subsequently converted into English words. Additionally, we compare the effectiveness of using katakana versus romaji as input characters. We develop 6 models of 2 types for our experiments: one with an English lexicon-filter, and the other without. For each type, we built 3 models, including a pair n-gram based on WFSTs and two sequence-to-sequence models leveraging LSTM and transformer. Our best performing model was the pair n-gram model with a lexicon-filter, directly transliterating from katakana to English.- Anthology ID:
- 2023.cawl-1.6
- Volume:
- Proceedings of the Workshop on Computation and Written Language (CAWL 2023)
- Month:
- July
- Year:
- 2023
- Address:
- Toronto, Canada
- Editors:
- Kyle Gorman, Richard Sproat, Brian Roark
- Venue:
- CAWL
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 43–49
- Language:
- URL:
- https://aclanthology.org/2023.cawl-1.6
- DOI:
- 10.18653/v1/2023.cawl-1.6
- Cite (ACL):
- Yuying Ren. 2023. Back-Transliteration of English Loanwords in Japanese. In Proceedings of the Workshop on Computation and Written Language (CAWL 2023), pages 43–49, Toronto, Canada. Association for Computational Linguistics.
- Cite (Informal):
- Back-Transliteration of English Loanwords in Japanese (Ren, CAWL 2023)
- PDF:
- https://preview.aclanthology.org/improve-issue-templates/2023.cawl-1.6.pdf