---
title: "Tokenization: A Survey for Modern NLP"
slug: tokenization-a-survey-for-modern-nlp
url: https://listedarticles.com/articles/tokenization-a-survey-for-modern-nlp
canonical_url: https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp
content_type: research
language: en
published_at: 2026-09-30T00:00:00.000Z
updated_at: 2026-10-02T03:18:44.731Z
author: "Marco Cognetta et al."
authored_by: human
publisher: "alphaXiv"
publisher_url: https://www.alphaxiv.org
topics: ["AI", "LLMs", "Research", "Machine Learning", "NLP", "Programming"]
license: all-rights-reserved
word_count: 376
reading_minutes: 2
citation: "Marco Cognetta et al., alphaXiv. \"Tokenization: A Survey for Modern NLP.\" 30 Sept 2026. https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Tokenization: A Survey for Modern NLP

> A comprehensive survey by 32+ tokenizer researchers covering subword algorithms, multilingual pitfalls, evaluation, theory, security, and alternatives—framing tokenization as a first-class bottleneck in the LLM era.

# Tokenization: A Survey for Modern NLP

*Marco Cognetta, Christopher Akiki, Pawan Sasanka Ammanamanchi, Catherine Arnett, Thomas Bauwens, and many co-authors (32+ tokenizer researchers) — submitted 30 Sept 2026 — [alphaXiv](https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp)*

## Abstract

While modern language models take raw text as their input and produce raw text as output, they do not operate over text directly. Hidden in the very first step of language model pipelines is tokenization, where raw text is converted to a sequence of tokens which the model is designed to process.

The dominant approach to tokenization is subword tokenization, which uses tokens whose granularity lies between that of characters (or bytes) and words. Subword tokenization is ubiquitous in modern natural language processing (NLP) due to its advantages: compared to word-level models, it solves the out-of-vocabulary problem while constraining the vocabulary size, reducing the number of embedding parameters and logits needed. Compared to character-level models, it improves the inductive bias of models by packing more information into a token while producing significantly shorter sequences.

While this technique is powerful, it is not without flaws. To name a few: it imposes a rigid system for how tokens are built, and one system may not generalize across languages; segmentation ambiguities create room for adversarial attacks; in multilingual settings, balancing the token allocation between high- and low-resource languages must be done explicitly; and the reliance on a static, pre-computed vocabulary means the tokenization process is disjoint from the model's training, preventing true end-to-end optimization. There is substantial room for innovation in this domain, yet progress is contingent upon a clear understanding of existing methodologies and current bottlenecks. This survey aims to serve that purpose.

We present tokenization in the context of the modern large language model (LLM) era and cover the breadth of areas that tokenization directly influences. We discuss algorithms, techniques, pitfalls, theory, and practice, along with jumping-off points for prospective researchers and practitioners. We hope that this survey brings more attention to the field of tokenization as an area in which much progress is still to be made.

## Scope (from authors’ announcement)

The survey covers algorithms, evaluations, multilinguality, encodings, theory, constrained generation, token healing, tokenizer security concerns, and alternatives (e.g., latent or visual tokenization). Sections include questions for prospective researchers and practitioners.

*Canonical page: https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp*
