Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Functional ultrasound imaging from scratch
Functional ultrasound imaging (fUSI) is a brain imaging modality that is just starting to be demonstrated in human studies. I’m sure many will be familiar with structural ultrasound imaging, which forms static images using ultrasound, as in a sonogram1. Functional ultrasound instead captures movies of brain tissue very rapidly, tracking subtle changes caused by the blood flow that follows neural activity.
11 min · 2,451 words
Build an LLM Tokenizer and Attention from Scratch in TypeScript
SitePoint tutorial that implements BPE tokenization, cosine similarity vector search, and scaled dot-product attention in TypeScript—exposing LLM primitives as ordinary readable code.
25 min · 5,827 words
Turn GLM-5.3-Flash into a Jev-like System One model
Johannes Hötter walks through turning GLM-5.3-Flash into a fast Jev-like “System One” decision model—typed options with probabilities in a single forward pass, matching Jev’s accuracy and speed.
11 min · 2,426 words
Teaching a World Model to Play Pokémon
Training a JEPA-style LeWorldModel on Pokémon Red screenshots and button presses, then using CEM latent planning to select a starter—why latent collapse, SIGReg, and rollout fine-tuning mattered, and 52/100 plans succeeding after tuning.
2 min · 574 words
Mixture of Experts (MoE) for Backend Engineers
A detailed visual guide to token routing, expert batching, weighted combination, and the memory and communication tradeoffs of MoE serving. Suppose a model has dozens of feed-forward subnetworks, but each token uses only two of them. The arithmetic per token can stay modest while the total weight set grows. Now place that model on eight GPUs. If every GPU stores all experts, memory can become the limit; if experts are split across GPUs, token activations must travel to whichever GPU owns their…
15 min · 3,515 words
V-JEPA: Learning Video Representations by Feature Prediction
A hands-on walkthrough of Meta’s V-JEPA: video tubelets, feature prediction, pretrained encoder features, temporal tests, and a small action-classification experiment.
25 min · 5,743 words
Heretic tutorial: automatic censorship removal for language models
A hands-on tutorial for Heretic, an open-source tool that automatically removes refusal/censorship behaviors from language models—setup, workflow, and what to watch for.
7 min · 1,688 words
Transformer Explainer: LLM Transformer Model Visually Explained
Georgia Tech’s Polo Club walks through GPT-2’s Transformer stack—embeddings, multi-head attention, MLP, sampling—with an interactive in-browser model for learning how next-token prediction works.
4 min · 827 words
Multimodal Agents: From Perception to Action
Illustrated notes from Berkeley’s LLM Agents lecture 7: OSWorld outcome checks, AgentTrek trajectories, TACO tools, and Aguvis grounding—why reading a screen is not the same as finishing the task.
8 min · 1,915 words
Training Search Agents with GRPO
Hands-on introduction to reinforcement learning by training a search agent with group-relative policy optimization (GRPO), with open rollouts, code, and reward-design lessons for LLM search.
28 min · 6,446 words
Gregory Gundersen reconstructs backpropagation from first principles, explaining why the algorithm needs a backward pass to compute neural network gradients efficiently.
6 min · 1,490 words