A detailed visual guide to token routing, expert batching, weighted combination, and the memory and communication tradeoffs of MoE serving.

Suppose a model has dozens of feed-forward subnetworks, but each token uses only two of them. The arithmetic per token can stay modest while the total weight set grows. Now place that model on eight GPUs. If every GPU stores all experts, memory can become the limit; if experts are split across GPUs, token activations must travel to whichever GPU owns their selected experts. What does that trade cost under a real request mix?

This article explains mixture of experts from first principles, then follows its execution across GPUs. We use explicitly constructed examples to show the mechanics. A dense model such as Llama 3.1 70B does not have routed expert layers; its weight and KV calculations must not be transferred to these MoE examples.