Tokenization: A Survey for Modern NLP

Marco Cognetta, Christopher Akiki, Pawan Sasanka Ammanamanchi, Catherine Arnett, Thomas Bauwens, and many co-authors (32+ tokenizer researchers) — submitted 30 Sept 2026 — alphaXiv

Abstract

While modern language models take raw text as their input and produce raw text as output, they do not operate over text directly. Hidden in the very first step of language model pipelines is tokenization, where raw text is converted to a sequence of tokens which the model is designed to process.

The dominant approach to tokenization is subword tokenization, which uses tokens whose granularity lies between that of characters (or bytes) and words. Subword tokenization is ubiquitous in modern natural language processing (NLP) due to its advantages: compared to word-level models, it solves the out-of-vocabulary problem while constraining the vocabulary size, reducing the number of embedding parameters and logits needed. Compared to character-level models, it improves the inductive bias of models by packing more information into a token while producing significantly shorter sequences.