Speculative Decoding with a Speculative Vocabulary
Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, Stylianos I. Venieris
Abstract
Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model comprising a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. While this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we show that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding method, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.- Anthology ID:
- 2026.findings-acl.2000
- Volume:
- Findings of the Association for Computational Linguistics: ACL 2026
- Month:
- July
- Year:
- 2026
- Address:
- San Diego, California, United States
- Editors:
- Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
- Venue:
- Findings
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 40240–40254
- Language:
- URL:
- https://preview.aclanthology.org/ingest-acl/2026.findings-acl.2000/
- DOI:
- Cite (ACL):
- Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, and Stylianos I. Venieris. 2026. Speculative Decoding with a Speculative Vocabulary. In Findings of the Association for Computational Linguistics: ACL 2026, pages 40240–40254, San Diego, California, United States. Association for Computational Linguistics.
- Cite (Informal):
- Speculative Decoding with a Speculative Vocabulary (Williams et al., Findings 2026)
- PDF:
- https://preview.aclanthology.org/ingest-acl/2026.findings-acl.2000.pdf