DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings

Yung-Sung Chuang; Rumen Dangovski; Hongyin Luo; Yang Zhang; Shiyu Chang; Marin Soljačić; Shang-Wen Li; Scott Yih; Yoon Kim; James Glass

doi:10.18653/v1/2022.naacl-main.311

DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings

Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James Glass

Abstract

We propose DiffCSE, an unsupervised contrastive learning framework for learning sentence embeddings. DiffCSE learns sentence embeddings that are sensitive to the difference between the original sentence and an edited sentence, where the edited sentence is obtained by stochastically masking out the original sentence and then sampling from a masked language model. We show that DiffSCE is an instance of equivariant contrastive learning, which generalizes contrastive learning and learns representations that are insensitive to certain types of augmentations and sensitive to other “harmful” types of augmentations. Our experiments show that DiffCSE achieves state-of-the-art results among unsupervised sentence representation learning methods, outperforming unsupervised SimCSE by 2.3 absolute points on semantic textual similarity tasks.

Anthology ID:: 2022.naacl-main.311
Volume:: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Month:: July
Year:: 2022
Address:: Seattle, United States
Venue:: NAACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 4207–4218
Language:
URL:: https://aclanthology.org/2022.naacl-main.311
DOI:: 10.18653/v1/2022.naacl-main.311
Bibkey:
Cite (ACL):: Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, and James Glass. 2022. DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4207–4218, Seattle, United States. Association for Computational Linguistics.
Cite (Informal):: DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings (Chuang et al., NAACL 2022)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingestion-script-update/2022.naacl-main.311.pdf
Video:: https://preview.aclanthology.org/ingestion-script-update/2022.naacl-main.311.mp4
Code: voidism/diffcse
Data: MPQA Opinion Corpus, MRPC, SICK, SST, STS Benchmark, Semantic Textual Similarity (2012 - 2016), SentEval

PDF Search Code Video