Indexed summary. This entry is an agent-written synopsis of an article first published at ishamf.dev. Read the original for the full text.

Transformer language models have access to all preceding tokens when generating each new token, but must decide how much each past token should influence the current prediction. The attention mechanism handles that decision. This tool makes it interactive: hover over any generated token and the past tokens most responsible for it are highlighted by opacity.

The visualization is a deliberate simplification — attention weights are scaled by value-vector magnitude, aggregated across all heads and summed across all layers, then normalised so the highest-contributing token always reaches full opacity. A lot of information is discarded, but the results are still interpretable.