Abstract
Code-switching is a commonly observed communicative phenomenon denoting a shift from one language to another within the same speech exchange. The analysis of code-switched data often becomes an assiduous task, owing to the limited availability of data. In this work, we propose converting code-switched data into its constituent high resource languages for exploiting both monolingual and cross-lingual settings. This conversion allows us to utilize the higher resource availability for its constituent languages for multiple downstream tasks. We perform experiments for two downstream tasks, sarcasm detection and hate speech detection in the English-Hindi code-switched setting. These experiments show an increase in 22% and 42.5% in F1-score for sarcasm detection and hate speech detection, respectively, compared to the state-of-the-art.- Anthology ID:
- 2020.aacl-srw.6
- Volume:
- Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop
- Month:
- December
- Year:
- 2020
- Address:
- Suzhou, China
- Venue:
- AACL
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 37–43
- Language:
- URL:
- https://aclanthology.org/2020.aacl-srw.6
- DOI:
- Cite (ACL):
- Kartikey Pant and Tanvi Dadu. 2020. Towards Code-switched Classification Exploiting Constituent Language Resources. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, pages 37–43, Suzhou, China. Association for Computational Linguistics.
- Cite (Informal):
- Towards Code-switched Classification Exploiting Constituent Language Resources (Pant & Dadu, AACL 2020)
- PDF:
- https://preview.aclanthology.org/ingestion-script-update/2020.aacl-srw.6.pdf