What Do Self-Supervised Speech Models Know About Words?

Ankita Pasad; Chung-Ming Chien; Shane Settle; Karen Livescu

doi:10.1162/tacl_a_00656

What Do Self-Supervised Speech Models Know About Words?

Ankita Pasad, Chung-Ming Chien, Shane Settle, Karen Livescu

Abstract

Many self-supervised speech models (S3Ms) have been introduced over the last few years, improving performance and data efficiency on various speech tasks. However, these empirical successes alone do not give a complete picture of what is learned during pre-training. Recent work has begun analyzing how S3Ms encode certain properties, such as phonetic and speaker information, but we still lack a proper understanding of knowledge encoded at the word level and beyond. In this work, we use lightweight analysis methods to study segment-level linguistic properties—word identity, boundaries, pronunciation, syntactic features, and semantic features—encoded in S3Ms. We present a comparative study of layer-wise representations from ten S3Ms and find that (i) the frame-level representations within each word segment are not all equally informative, and (ii) the pre-training objective and model size heavily influence the accessibility and distribution of linguistic information across layers. We also find that on several tasks—word discrimination, word segmentation, and semantic sentence similarity—S3Ms trained with visual grounding outperform their speech-only counterparts. Finally, our task-based analyses demonstrate improved performance on word segmentation and acoustic word discrimination while using simpler methods than prior work.1

Anthology ID:: 2024.tacl-1.21
Volume:: Transactions of the Association for Computational Linguistics, Volume 12
Month:
Year:: 2024
Address:: Cambridge, MA
Venue:: TACL
SIG:
Publisher:: MIT Press
Note:
Pages:: 372–391
Language:
URL:: https://preview.aclanthology.org/fix-sig-urls/2024.tacl-1.21/
DOI:: 10.1162/tacl_a_00656
Bibkey:
Cite (ACL):: Ankita Pasad, Chung-Ming Chien, Shane Settle, and Karen Livescu. 2024. What Do Self-Supervised Speech Models Know About Words?. Transactions of the Association for Computational Linguistics, 12:372–391.
Cite (Informal):: What Do Self-Supervised Speech Models Know About Words? (Pasad et al., TACL 2024)
Copy Citation:
PDF:: https://preview.aclanthology.org/fix-sig-urls/2024.tacl-1.21.pdf

PDF Cite Search Fix data