Abstract
Text transcripts without punctuation or sentence boundaries are hard to comprehend for both humans and machines. Punctuation marks play a vital role by providing meaning to the sentence and incorrect use or placement of punctuation marks can often alter it. This can impact downstream tasks such as language translation and understanding, pronoun resolution, text summarization, etc. for humans and machines. An automated punctuation restoration (APR) system with minimal human intervention can improve comprehension of text and help users write better. In this paper we describe a multitask modeling approach as a system to restore punctuation in multiple high resource – Germanic (English and German), Romanic (French)– and low resource languages – Indo-Aryan (Hindi) Dravidian (Tamil) – that does not require extensive knowledge of grammar or syntax of a given language for both spoken and written form of text. For German language and the given Indic based languages this is the first towards restoring punctuation and can serve as a baseline for future work.- Anthology ID:
- 2021.eacl-demos.37
- Volume:
- Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations
- Month:
- April
- Year:
- 2021
- Address:
- Online
- Editors:
- Dimitra Gkatzia, Djamé Seddah
- Venue:
- EACL
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 312–320
- Language:
- URL:
- https://aclanthology.org/2021.eacl-demos.37
- DOI:
- 10.18653/v1/2021.eacl-demos.37
- Cite (ACL):
- Varnith Chordia. 2021. PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pages 312–320, Online. Association for Computational Linguistics.
- Cite (Informal):
- PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text (Chordia, EACL 2021)
- PDF:
- https://preview.aclanthology.org/improve-issue-templates/2021.eacl-demos.37.pdf
- Code
- VarnithChordia/Multlingual_Punctuation_restoration