EUREKA: EUphemism Recognition Enhanced through Knn-based methods and Augmentation

Sedrick Scott Keh, Rohit Bharadwaj, Emmy Liu, Simone Tedeschi, Varun Gangal, Roberto Navigli


Abstract
We introduce EUREKA, an ensemble-based approach for performing automatic euphemism detection. We (1) identify and correct potentially mislabelled rows in the dataset, (2) curate an expanded corpus called EuphAug, (3) leverage model representations of Potentially Euphemistic Terms (PETs), and (4) explore using representations of semantically close sentences to aid in classification. Using our augmented dataset and kNN-based methods, EUREKA was able to achieve state-of-the-art results on the public leaderboard of the Euphemism Detection Shared Task, ranking first with a macro F1 score of 0.881.
Anthology ID:
2022.flp-1.15
Volume:
Proceedings of the 3rd Workshop on Figurative Language Processing (FLP)
Month:
December
Year:
2022
Address:
Abu Dhabi, United Arab Emirates (Hybrid)
Venue:
FLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
111–117
Language:
URL:
https://aclanthology.org/2022.flp-1.15
DOI:
Bibkey:
Cite (ACL):
Sedrick Scott Keh, Rohit Bharadwaj, Emmy Liu, Simone Tedeschi, Varun Gangal, and Roberto Navigli. 2022. EUREKA: EUphemism Recognition Enhanced through Knn-based methods and Augmentation. In Proceedings of the 3rd Workshop on Figurative Language Processing (FLP), pages 111–117, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
Cite (Informal):
EUREKA: EUphemism Recognition Enhanced through Knn-based methods and Augmentation (Keh et al., FLP 2022)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingestion-script-update/2022.flp-1.15.pdf