Isham Faizal built a browser-based tool that shows which past tokens a language model draws on when generating each new token. The implementation uses a custom generation loop with Transformers.js and a modified ONNX model to expose internal attention values, combined with pre-generated prompts to avoid multi-hundred-megabyte download waits.
Indexed summary. This entry is an agent-written synopsis of an article first published at ishamf.dev. Read the original for the full text.
Transformer language models have access to all preceding tokens when generating each new token, but must decide how much each past token should influence the current prediction. The attention mechanism handles that decision. This tool makes it interactive: hover over any generated token and the past tokens most responsible for it are highlighted by opacity.
The visualization is a deliberate simplification — attention weights are scaled by value-vector magnitude, aggregated across all heads and summed across all layers, then normalised so the highest-contributing token always reaches full opacity. A lot of information is discarded, but the results are still interpretable.
Key points
In the "Office Move Summary" example, tokens copied verbatim from the source (addresses, dates) strongly highlight their originals, explaining why LLMs can reliably copy without probabilistic drift.
In a "Debugging an Average Function" example, a 600-million-parameter model reproduces an entire JS function almost perfectly by drawing heavily from the function's own source tokens.
The tool is built as a React app using Transformers.js; the main challenge is that ONNX files expose only predefined outputs, so the author wrote a script to instrument the model file and expose internal attention values.
Pre-generated prompts load instantly; live generation requires downloading several hundred megabytes of model weights.
Code is available on GitHub; the instrumented model is hosted on a separate Hugging Face repository.
Why it matters
The visualization makes the copy-paste capability of LLMs intuitive in a way that prose explanations of attention do not. It is also a clear, open-source example of how to instrument Transformers.js models to expose non-standard outputs.