Transformer is a neural network architecture that has fundamentally changed the approach to Artificial Intelligence. Transformer was first introduced in the seminal paper "Attention is All You Need" in 2017 and has since become the go-to architecture for deep learning models, powering text-generative models like OpenAI's GPT, Meta's Llama, and Google's Gemini. Beyond text, Transformer is also applied in audio generation, image recognition, protein structure prediction, and even game playing, demonstrating its versatility across numerous domains.
Fundamentally, text-generative Transformer models operate on the principle of next-token prediction: given a text prompt from the user, what is the most probable next token (a word or part of a word) that will follow this input? The core innovation and power of Transformers lie in their use of self-attention mechanism, which allows them to process entire sequences and capture long-range dependencies more effectively than previous architectures.
GPT-2 family of models are prominent examples of text-generative Transformers. Transformer Explainer is powered by the GPT-2 (small) model which has 124 million parameters. While it is not the latest or most powerful Transformer model, it shares many of the same architectural components and principles found in the current state-of-the-art models making it an ideal starting point for understanding the basics.
Every text-generative Transformer consists of these three key components:
- Embedding: Text input is divided into smaller units called tokens, which can be words or subwords. These tokens are converted into numerical vectors called embeddings, which capture the semantic meaning of words.
- Transformer Block is the fundamental building block of the model that processes and transforms the input data. Each block includes:
- Attention Mechanism, the core component of the Transformer block. It allows tokens to communicate with other tokens, capturing contextual information and relationships between words.
- MLP (Multilayer Perceptron) Layer, a feed-forward network that operates on each token independently. While the goal of the attention layer is to route information between tokens, the goal of the MLP is to refine each token's representation.
- Output Probabilities: The final linear and softmax layers transform the processed embeddings into probabilities, enabling the model to make predictions about the next token in a sequence.
Embedding
To convert a prompt into embedding, we need to (1) tokenize the input, (2) obtain token embeddings, (3) add positional information, and finally (4) add up token and position encodings to get the final embedding.
Step 1: Tokenization
Tokenization is the process of breaking down the input text into smaller, more manageable pieces called tokens. These tokens can be a word or a subword. The full vocabulary of tokens is decided before training the model: GPT-2's vocabulary has 50,257 unique tokens.
Step 2. Token Embedding
GPT-2 (small) represents each token in the vocabulary as a 768-dimensional vector. These embedding vectors are stored in a matrix of shape (50,257, 768), containing approximately 39 million parameters.
Step 3. Positional Encoding
The Embedding layer also encodes information about each token's position in the input prompt. GPT-2 trains its own positional encoding matrix from scratch.
Step 4. Final Embedding
Finally, we sum the token and positional encodings to get the final embedding representation. This combined representation captures both the semantic meaning of the tokens and their position in the input sequence.
The core of the Transformer's processing lies in the Transformer block, which comprises multi-head self-attention and a Multi-Layer Perceptron layer. The GPT-2 (small) model consists of 12 such blocks.
Multi-Head Self-Attention
The self-attention mechanism enables the model to capture relationships among tokens in a sequence. Multiple attention heads allow the model to consider these relationships from different perspectives.
Query, Key, and Value Matrices
Each token's embedding vector is transformed into three vectors: Query (Q), Key (K), and Value (V). A useful analogy: Query is the search text; Key is the title of each result; Value is the content of matched pages.
Multi-Head Splitting, Masked Self-Attention, Output
Query, Key, and Value vectors are split into multiple heads—in GPT-2 (small)'s case, into 12 heads. Masked self-attention prevents access to future tokens. Softmax converts scores into probabilities. Head outputs are concatenated and linearly projected.
MLP: Multi-Layer Perceptron
After multi-head self-attention, concatenated outputs pass through an MLP: a first linear layer expands dimensionality four-fold from 768 to 3072 with GELU, then a second linear layer compresses back to 768.
Output Probabilities
After all Transformer blocks, a final linear layer projects into a 50,257-dimensional logit space over the vocabulary; softmax yields next-token probabilities. Temperature, top-k, and top-p sampling control determinism vs. diversity.
Auxiliary Architectural Features
Layer Normalization, Dropout, and Residual Connections stabilize training, prevent overfitting, and mitigate vanishing gradients.
Interactive Features
Transformer Explainer runs a live GPT-2 (small) model in the browser (nanoGPT → ONNX Runtime), with a Svelte + D3.js interface for attention maps, temperature, and sampling controls.
Created by Aeree Cho, Grace C. Kim, Alexander Karpekov, Alec Helbling, Jay Wang, Seongmin Lee, Benjamin Hoover, and Polo Chau at the Georgia Institute of Technology.