Publisher

Cohere

Support automation that's actually intelligent

Showing 1–1 of 1 article

  • Cohere's North Mini Code Megakernel Serving Engine

    Today, Cohere presents a serving engine for North Mini Code built around a decode megakernel: BF16 on a single H100, 1.25× - 1.41× faster than vLLM end-to-end. Explore the code behind the serving engine on GitHub. Most LLM serving stacks still treat each forward pass as a sequence of kernels: launch QKV, wait; launch attention, wait; launch the MoE, wait. Each launch is fine on its own. The problem is the waiting in between. At small batch sizes, the GPU spends a surprising fraction of every decode step waiting rather than computing.

    Blog post · AI · LLMs · Performance · Machine Learning

    24 min · 5,534 words