Sander Land

2026

Which Pieces Does Unigram Tokenization Really Need?
Sander Land | Yuval Pinter
Findings of the Association for Computational Linguistics: ACL 2026

The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegance, its implementation in practice is complex, limiting its adoption to the SentencePiece package and adapters thereof. We bridge this gap between theory and practice by providing a clear guide to implementation and parameter choices. We also identify a simpler algorithm that accepts slightly higher training loss in exchange for improved compression.

2024

pdf bib abs

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Sander Land | Max Bartolo
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous SolidGoldMagikarp token, to induce unwanted model behaviour. Although such ‘glitch tokens’, tokens present in the tokenizer vocabulary but that are nearly or entirely absent during model training, have been observed across various models, a reliable method to identify and address them has been missing. We present a comprehensive analysis of Large Language Model tokenizers, specifically targeting this issue of detecting under-trained tokens. Through a combination of tokenizer analysis, model weight-based indicators, and prompting techniques, we develop novel and effective methods for automatically detecting these problematic tokens. Our findings demonstrate the prevalence of such tokens across a diverse set of models and provide insights into improving the efficiency and safety of language models.

Co-authors

Max Bartolo 1
Yuval Pinter 1

Venues

EMNLP1
Findings1

Fix author