{"article":{"slug":"the-rise-of-overfit-inference-engines","title":"The Rise of Overfit Inference Engines","subtitle":null,"summary":"A hands-on essay on how local LLM inference stacks are overfit to specific models and hardware—and what that means for anyone chasing reproducible homelab performance with llama.cpp and friends.","content_type":"essay","language":"en","canonical_url":"https://carteakey.dev/blog/local-inference/the-rise-of-overfit-inference-engines/","author":{"name":"carteakey","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"carteakey.dev","url":null,"listing_slug":null,"listing":null},"topics":[{"name":"llms","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"machine-learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"hardware","slug":"hardware","url":"https://listedarticles.com/topics/hardware"},{"name":"engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2324,"reading_minutes":10,"published_at":"2026-09-30T12:00:00.000Z","added_at":"2026-10-03T20:17:01.654Z","updated_at":"2026-10-03T20:17:01.654Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-rise-of-overfit-inference-engines","markdown_url":"https://listedarticles.com/articles/the-rise-of-overfit-inference-engines.md","example":false,"citation":"carteakey, carteakey.dev. \"The Rise of Overfit Inference Engines.\" 30 Sept 2026. https://carteakey.dev/blog/local-inference/the-rise-of-overfit-inference-engines/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://carteakey.dev/blog/local-inference/the-rise-of-overfit-inference-engines/"},"body_markdown":"{% image_cc \"./src/static/img/local-inference/overfit-inference-midwit-meme.webp\", \"Bell-curve meme. The low-IQ end and the hooded high-IQ end both say: I write inference code for my specific model and hardware. The crying midwit in the middle says: I write inference code that generalizes across all models and all hardware.\", \"w-full border border-surface-border my-6\", \"The whole post, in one meme.\" %}\n\n## The surprise\n\nFor a week I'd been tuning [Qwen3.8-Flash-Next](/blog/local-inference/running-qwen3-8-flash-next-locally/), a 125B-parameter MoE, on {% device \"yeti-cachy\" %} (my main homelab node: an i5-12600K with 64 GB DDR5 and an RTX 4070 12GB I paid $500 for). After a lot of flags and two experimental branches, `llama.cpp` reached **27 tok/s**. For an offloaded 125B model on a 12 GB card, that felt like the ceiling.\n\nThen I tried **Strata**, a runtime that supports exactly this model on NVIDIA consumer cards.\n\n```text\nstrata serve: prompt 59787 tokens = 0 reused + 59787 read in 29701 ms (2013.0 tok/s), 512 generated in 9631 ms (53.2 tok/s), drafts accepted 244 of 331\n```\n\nSame box, 60,000 tokens of context: **53 tok/s**. About twice my best llama.cpp number, with no hardware changes.\n\nMy unfiltered reaction was: *what in the fuck?*\n\n{% diagram_card {\n  src: \"./src/static/img/diagrams/qwen38-flash-next-throughput-progression.png\",\n  alt: \"Line chart of Qwen3.8-Flash-Next decode throughput on an RTX 4070 12GB in tokens per second across nine steps: Powersave 6.5, CPU governor 12.2, SSD mmap 15.2, q8 KV plus fit 18.9, master 19.35, MTP V1 20.65, MTP V2 27.06, Strata dynamic cache 60.3, Strata peak 90.2\",\n  kicker: \"Decode · RTX 4070 12GB + 64GB DDR5\",\n  title: \"Seven Tuning Steps, Then One Runtime Swap\",\n  badge: \"27.06 → 53.2 tok/s at 60k\",\n  caption: \"The first seven points are tuning steps on llama.cpp, not separate runtimes. 60.3 is Strata's best short-prompt row; at a 60k context it holds 53.2. 90.2 is a warm-draft burst on repetitive code, not sustained throughput.\"\n} %}\n\nQuants and settings differ between the two engines, so that's a comparison of complete serving setups, not a controlled experiment. The measurements, memory layout and caveats are in the technical write-up, [Strata on an RTX 4070](/blog/local-inference/strata-on-an-rtx-4070/). This post is about what I think it means.\n\nMy bet: for power users on fixed hardware, **disposable, overfit engines are going to beat general-purpose runtimes on speed**, and general runtimes will keep the portability crown.\n\n## The new breed\n\nStrata isn't alone. In the past few months a handful of narrow inference engines have appeared, each built for a short list of models on one hardware family, and each beating general runtimes inside its niche.\n\n{% wide %}\n| Engine | Target hardware | Models | Notable trick |\n| --- | --- | --- | --- |\n| [Strata](https://github.com/Niko1221/Strata) | Consumer NVIDIA RTX, 12 GB+ | Qwen3.8-Flash-Next | Per-expert VRAM cache across all layers, native MTP |\n| [ninfer](https://github.com/Neroued/ninfer) | One RTX 5090 (community forks for 3090) | A closed list of Qwen checkpoints | Scratch-written C++/CUDA, no offloading, no multi-GPU |\n| [DwarfStar](https://github.com/antirez/ds4) (`ds4`) | Metal, CUDA, ROCm; two Macs over RDMA | DeepSeek V4/V4.1 Flash and V4 Pro first, plus GLM 5.x and Qwen3.8-Flash-Next | Aggressive routed-expert quants, compressed KV cache on SSD |\n| [Splash](https://github.com/incoai/splash) | Apple Silicon M3+ | Qwen family | DFlash 2 speculation, per-model kernels and memory plans |\n| [llamAmpere](https://github.com/JakeATX/llamAmpere) | RTX 3090 / 3090 Ti (SM86) | Mainly Qwen3.8-27B | llama.cpp fork with TurboQuant KV and custom verify kernels |\n| [gufo](https://github.com/gufo-org/gufo) | AMD Strix Halo (Ryzen AI MAX+ 395) | A small curated set | Custom HIP kernels, DFlash2 + MTP, continuous batching |\n{% endwide %}\n\nBy normal software standards, these can look like bad codebases: tightly coupled, hard to port, built around narrow assumptions. In inference, those trade-offs are part of why they're fast. They shape the compute graph, memory layout and kernels around their chosen models and hardware.\n\n## A prediction you can check\n\nMost of these will be abandoned soon, and that's fine.\n\nWhen the next model generation changes its routing scheme or attention layout, an engine built around the old one won't refactor. It'll stall, and a new one-off will show up for the new model within days or weeks.\n\nTo make that testable: **by April 2027, at least four of the six engines above will have gone 60 days without a commit to their default branch**. I'll count upstream commits, not activity in forks, and I'll check back and report either way.\n\n## Why now: three reasons\n\n### 1. Generality is a performance tax\n\nThe more a codebase supports, the harder it is to make changes that break assumptions. `llama.cpp` is one of the best pieces of open-source engineering around, and it runs dozens of model families across CUDA, ROCm, Metal, Vulkan, SYCL and a long list of CPU architectures.\n\nThat breadth has a cost. A change to how experts move across PCIe can't break Metal's unified memory path. A CUDA optimization can't break older cards. A per-expert VRAM cache like Strata's has to be threaded through the shared graph scheduler and allocator that every backend depends on. Each of those is solvable, but slowly, and through review.\n\nA small codebase with one target doesn't have those constraints. It can make the radical change on a Tuesday.\n\n### 2. AI coding made engines cheap to write\n\nA high-performance CUDA runtime used to need a senior systems engineer and months of profiling. Today a capable developer with a coding agent can write, benchmark and fix kernels far faster. A fixed model and machine make that work easier to scope: hand the agent one tensor layout or one fusion, then measure.\n\nThere's at least one public data point. [DwarfStar's README](https://github.com/antirez/ds4) says it's developed \"with strong assistance from AI coding agents,\" with humans leading the ideas, testing and debugging. I don't know how much of Strata was agent-written. My guess is that the cost of a purpose-built runtime has dropped from months of specialist time to weeks or less, but I don't have hard numbers.\n\n### 3. \"Make tok/s go up\" is an unusually crisp objective\n\nMost software resists automation because the goal is fuzzy. Inference optimization is closer to machine-checkable:\n\n1. **Input:** a fixed weights file and token IDs.\n2. **Constraint:** outputs stay within a tolerance of a reference, for example greedy-token agreement or low KL divergence on reference logits.\n3. **Objective:** more tokens per second in less memory.\n\nYou can point an agent loop at that and let it iterate against hardware counters.\n\nIt isn't fully specified, though. Prefill latency, power, concurrency, context length, sampling behavior and reliability all matter, and a loose tolerance invites the loop to trade away quality it doesn't measure, such as long-context accuracy or rare tokens. The tighter the parity check, the more you can trust the result. That's also why the quality check is the most important open item in my Strata write-up.\n\n## The twist: general engines become the parts bin\n\nIt would be easy to frame this as napkin runtimes versus general runtimes. The engines themselves suggest something messier.\n\nDwarfStar borrows kernels and quant formats from llama.cpp's GGML. llamAmpere is a llama.cpp fork. The GGUF format, the quant types and many of the kernels these projects start from came out of years of general-purpose work.\n\nSo the more likely future isn't one replacing the other. General engines keep doing the slow, broad work: formats, quantization research, correct reference kernels for every backend. Napkin runtimes take those parts, bolt them into something overfit for one model and one card, and throw the assembly away when the next model arrives.\n\nGeneral engines become the parts bin. Napkin runtimes are what power users build from it.\n\n## The console analogy\n\nGame developers already know this pattern. A console studio targets one exact chip, one memory bus and one cache hierarchy, and that fixed target gives it permission to specialize aggressively. PC developers can't, because their engine has to run on whatever hardware shows up, so they pay a generality tax in abstraction layers and defensive code paths.\n\n{% image_cc \"./src/static/img/diagrams/overfit-console-vs-pc.png\", \"Two-column comparison. PC Model, Generality Tax: a stack of llama.cpp / vLLM / Ollama, Graph schedulers, Multi-OS, CUDA / ROCm / CPU and 50+ model families. Console Model, Direct-to-Metal: Strata / ninfer / Splash pointing straight at Ada SM89 / PCIe 4.0.\", \"sketch-draw\", \"PC model vs console model of inference runtimes. A generalist runtime pays for graph schedulers, multi-OS and multi-vendor support, and dozens of model families. An overfit runtime is tuned for one target and drops everything else.\" %}\n\nFor years nobody wrote console-style engines for a single PC configuration, because no one could justify the engineering cost for an audience of one card. Cheaper engine-writing changes that math. If a runtime for one model on one GPU costs weeks instead of years, your desktop can be treated like a console: a fixed target worth specializing for.\n\n## What survives the churn\n\nIf runtimes are disposable, the layers around them have to be stable. Nobody wants a new CLI, UI and SDK every time an engine dies.\n\n### The HTTP API\n\nThe OpenAI Chat Completions API and the Anthropic Messages API are what clients speak. As long as a napkin runtime exposes one of them, it plugs into Cursor, [Cline](https://github.com/cline/cline), Open WebUI, Aider and the rest without anyone noticing what's underneath.\n\n### Model routers\n\nRunning several one-off engines means juggling processes and ports. That's what [`llama-swap`](https://github.com/mostlygeek/llama-swap) and my [L3MS](https://github.com/carteakey/l3ms) supervisor are for. My current tiers:\n\n| Model name | Engine | Role | Decode |\n| --- | --- | --- | --- |\n| `qwen38-flash-next-plat` | Strata v0.1.30 | Fast default, 128k context | ~53–60 tok/s |\n| `qwen38-flash-next-plat-vision` | Strata v0.1.30 + CPU vision encoder | Fast image tasks | Untested with images; the CPU encoder mostly adds time to first token |\n| `qwen38-flash-next` | llama.cpp master | Stable fallback, 96k context | 20.8 tok/s |\n| `qwen38-flash-next-vision` | llama.cpp master + mmproj | Image tasks fallback, 16k context | 18.2–18.6 tok/s |\n| `qwen38-flash-next-mtp` | llama.cpp MTP branch | Experimental | 25.3–27.1 tok/s |\n| `gemma-4-26b-qat-mtp` | llama.cpp + MTP assistant | Smaller MoE, mostly on the GPU | [~100 tok/s](/blog/local-inference/gemma-4-26b-qat-mtp/) |\n\nllama-swap listens on one port. When a client asks for a model, it stops whatever was loaded, starts the right engine with its flags, proxies the request, and shuts it down after 10 idle minutes. The client sees one endpoint; engines swap underneath like cartridges. Trying Strata didn't mean replacing my workflow, and llama.cpp is still one model name away when the experiment breaks.\n\n### Hardware-specific communities (speculative)\n\nI'd expect tuning groups to form around exact hardware combinations rather than broad forums. Think a community for a 24 GB card with 64 GB of DDR5, or for one Mac memory tier, trading engine builds, expert routing profiles and placement recipes for their exact setup. This is extrapolation; I haven't seen it happen yet.\n\n## The cost: trust\n\nThe same properties that make napkin runtimes fast make them risky to run. They're young, maintained by one person or a small group, change daily, and ship native code that pins memory and talks straight to the GPU driver. Few people review them, and nobody issues security advisories for a project that will be abandoned in six months.\n\nIf this becomes the normal way to run local models, power users take on a supply-chain problem they didn't have with a well-reviewed general runtime. A source build and a pinned revision make the work inspectable; they don't make it audited. The minimum I'd suggest: build from source, pin a commit, skim the diff before updating, and don't run them on machines that hold anything you can't afford to lose. Hardware communities that share vetted builds could help here, or make it worse.\n\n## How general engines could win back the lead\n\nI don't think general engines go away. Three ways they could close the gap:\n\n**Recipes as plugins.** Alongside a model file, you download a recipe for your hardware: fused kernels, an expert placement mask and a speculation schedule built for, say, an RTX 4070 with DDR5. The general engine becomes a thin, trusted host that loads it. This is the parts-bin idea run in reverse, and it would also ease the trust problem, since the host stays reviewed. It's a proposed design; I'm not aware of a project that ships this today.\n\n**Agents maintaining the variants upstream.** The same agents that write one-off engines could generate and test hardware-specialized kernel variants inside a general repo, for every GPU generation, without maintainers hand-writing each one. The bottleneck shifts from writing code to reviewing it.\n\n**Compilers finally deliver.** If Mojo/MAX, Triton or MLIR-based stacks get good enough to take a high-level model graph and a hardware description and emit near-optimal code automatically, hand-overfit engines lose their edge overnight. Until then, a person or agent writing for one exact GPU will usually beat compiler output.\n\n## Don't marry your inference engine\n\nFor two years, running local models meant one playbook: install llama.cpp or vLLM, pick a context size, offload what fits. That playbook still works, and for most people it's still the right one.\n\nBut when a runtime written for one model and one GPU doubles your best decode speed and holds it flat across 60,000 tokens, the generality tax gets hard to ignore. So keep your serving layer stable, treat the engine underneath as replaceable, and don't expect this year's fastest runtime to support next year's model.\n\nNapkin software: cheap to write, very good at the meal in front of you, and fine to throw away when the next course arrives.\n\n{% callout \"note\", \"What this argument rests on\" %}\nOne model and one machine that I measured myself, with a missing ablation and no quality A/B yet, plus five other projects whose claims I'm taking at face value. The trend is real enough to notice. Whether it lasts is what the April 2027 check is for. I'd like to hear where you think it breaks.\n{% endcallout %}\n\n## Changelog\n\n| Date | Note |\n| --- | --- |\n| 2026-10-03 | Moved the benchmarks, memory layout and reproduction details to a separate [Strata post](/blog/local-inference/strata-on-an-rtx-4070/). Reframed the headline speedup against the best llama.cpp setup (about 2×), added a checkable prediction, the parts-bin section and a trust section, softened the objective-function claim, trimmed the console analogy, updated the DwarfStar description, and dropped an unverified reference. |\n| 2026-10-01 | Corrected the bandwidth arithmetic and removed the double-counted cache and MTP gain, aligned the hardware specs, linked all six engines, hedged the unsourced claims, moved the IQ3_S and NAS-archive notes to the Flash-Next post, and tightened the prose. |\n| 2026-09-30 | Initial post. |","body_html":"<p>{% image_cc &quot;./src/static/img/local-inference/overfit-inference-midwit-meme.webp&quot;, &quot;Bell-curve meme. The low-IQ end and the hooded high-IQ end both say: I write inference code for my specific model and hardware. The crying midwit in the middle says: I write inference code that generalizes across all models and all hardware.&quot;, &quot;w-full border border-surface-border my-6&quot;, &quot;The whole post, in one meme.&quot; %}</p>\n<h2 id=\"the-surprise\">The surprise</h2>\n<p>For a week I&#39;d been tuning <a href=\"/blog/local-inference/running-qwen3-8-flash-next-locally/\">Qwen3.8-Flash-Next</a>, a 125B-parameter MoE, on {% device &quot;yeti-cachy&quot; %} (my main homelab node: an i5-12600K with 64 GB DDR5 and an RTX 4070 12GB I paid $500 for). After a lot of flags and two experimental branches, <code>llama.cpp</code> reached <strong>27 tok/s</strong>. For an offloaded 125B model on a 12 GB card, that felt like the ceiling.</p>\n<p>Then I tried <strong>Strata</strong>, a runtime that supports exactly this model on NVIDIA consumer cards.</p>\n<pre><code class=\"language-text\">strata serve: prompt 59787 tokens = 0 reused + 59787 read in 29701 ms (2013.0 tok/s), 512 generated in 9631 ms (53.2 tok/s), drafts accepted 244 of 331</code></pre>\n<p>Same box, 60,000 tokens of context: <strong>53 tok/s</strong>. About twice my best llama.cpp number, with no hardware changes.</p>\n<p>My unfiltered reaction was: <em>what in the fuck?</em></p>\n<p>{% diagram_card {\n  src: &quot;./src/static/img/diagrams/qwen38-flash-next-throughput-progression.png&quot;,\n  alt: &quot;Line chart of Qwen3.8-Flash-Next decode throughput on an RTX 4070 12GB in tokens per second across nine steps: Powersave 6.5, CPU governor 12.2, SSD mmap 15.2, q8 KV plus fit 18.9, master 19.35, MTP V1 20.65, MTP V2 27.06, Strata dynamic cache 60.3, Strata peak 90.2&quot;,\n  kicker: &quot;Decode · RTX 4070 12GB + 64GB DDR5&quot;,\n  title: &quot;Seven Tuning Steps, Then One Runtime Swap&quot;,\n  badge: &quot;27.06 → 53.2 tok/s at 60k&quot;,\n  caption: &quot;The first seven points are tuning steps on llama.cpp, not separate runtimes. 60.3 is Strata&#39;s best short-prompt row; at a 60k context it holds 53.2. 90.2 is a warm-draft burst on repetitive code, not sustained throughput.&quot;\n} %}</p>\n<p>Quants and settings differ between the two engines, so that&#39;s a comparison of complete serving setups, not a controlled experiment. The measurements, memory layout and caveats are in the technical write-up, <a href=\"/blog/local-inference/strata-on-an-rtx-4070/\">Strata on an RTX 4070</a>. This post is about what I think it means.</p>\n<p>My bet: for power users on fixed hardware, <strong>disposable, overfit engines are going to beat general-purpose runtimes on speed</strong>, and general runtimes will keep the portability crown.</p>\n<h2 id=\"the-new-breed\">The new breed</h2>\n<p>Strata isn&#39;t alone. In the past few months a handful of narrow inference engines have appeared, each built for a short list of models on one hardware family, and each beating general runtimes inside its niche.</p>\n<p>{% wide %}</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Engine</th><th>Target hardware</th><th>Models</th><th>Notable trick</th></tr></thead><tbody><tr><td><a href=\"https://github.com/Niko1221/Strata\" rel=\"nofollow ugc noopener\">Strata</a></td><td>Consumer NVIDIA RTX, 12 GB+</td><td>Qwen3.8-Flash-Next</td><td>Per-expert VRAM cache across all layers, native MTP</td></tr><tr><td><a href=\"https://github.com/Neroued/ninfer\" rel=\"nofollow ugc noopener\">ninfer</a></td><td>One RTX 5090 (community forks for 3090)</td><td>A closed list of Qwen checkpoints</td><td>Scratch-written C++/CUDA, no offloading, no multi-GPU</td></tr><tr><td><a href=\"https://github.com/antirez/ds4\" rel=\"nofollow ugc noopener\">DwarfStar</a> (<code>ds4</code>)</td><td>Metal, CUDA, ROCm; two Macs over RDMA</td><td>DeepSeek V4/V4.1 Flash and V4 Pro first, plus GLM 5.x and Qwen3.8-Flash-Next</td><td>Aggressive routed-expert quants, compressed KV cache on SSD</td></tr><tr><td><a href=\"https://github.com/incoai/splash\" rel=\"nofollow ugc noopener\">Splash</a></td><td>Apple Silicon M3+</td><td>Qwen family</td><td>DFlash 2 speculation, per-model kernels and memory plans</td></tr><tr><td><a href=\"https://github.com/JakeATX/llamAmpere\" rel=\"nofollow ugc noopener\">llamAmpere</a></td><td>RTX 3090 / 3090 Ti (SM86)</td><td>Mainly Qwen3.8-27B</td><td>llama.cpp fork with TurboQuant KV and custom verify kernels</td></tr><tr><td><a href=\"https://github.com/gufo-org/gufo\" rel=\"nofollow ugc noopener\">gufo</a></td><td>AMD Strix Halo (Ryzen AI MAX+ 395)</td><td>A small curated set</td><td>Custom HIP kernels, DFlash2 + MTP, continuous batching</td></tr></tbody></table></div>\n<p>{% endwide %}</p>\n<p>By normal software standards, these can look like bad codebases: tightly coupled, hard to port, built around narrow assumptions. In inference, those trade-offs are part of why they&#39;re fast. They shape the compute graph, memory layout and kernels around their chosen models and hardware.</p>\n<h2 id=\"a-prediction-you-can-check\">A prediction you can check</h2>\n<p>Most of these will be abandoned soon, and that&#39;s fine.</p>\n<p>When the next model generation changes its routing scheme or attention layout, an engine built around the old one won&#39;t refactor. It&#39;ll stall, and a new one-off will show up for the new model within days or weeks.</p>\n<p>To make that testable: <strong>by April 2027, at least four of the six engines above will have gone 60 days without a commit to their default branch</strong>. I&#39;ll count upstream commits, not activity in forks, and I&#39;ll check back and report either way.</p>\n<h2 id=\"why-now-three-reasons\">Why now: three reasons</h2>\n<h3 id=\"1-generality-is-a-performance-tax\">1. Generality is a performance tax</h3>\n<p>The more a codebase supports, the harder it is to make changes that break assumptions. <code>llama.cpp</code> is one of the best pieces of open-source engineering around, and it runs dozens of model families across CUDA, ROCm, Metal, Vulkan, SYCL and a long list of CPU architectures.</p>\n<p>That breadth has a cost. A change to how experts move across PCIe can&#39;t break Metal&#39;s unified memory path. A CUDA optimization can&#39;t break older cards. A per-expert VRAM cache like Strata&#39;s has to be threaded through the shared graph scheduler and allocator that every backend depends on. Each of those is solvable, but slowly, and through review.</p>\n<p>A small codebase with one target doesn&#39;t have those constraints. It can make the radical change on a Tuesday.</p>\n<h3 id=\"2-ai-coding-made-engines-cheap-to-write\">2. AI coding made engines cheap to write</h3>\n<p>A high-performance CUDA runtime used to need a senior systems engineer and months of profiling. Today a capable developer with a coding agent can write, benchmark and fix kernels far faster. A fixed model and machine make that work easier to scope: hand the agent one tensor layout or one fusion, then measure.</p>\n<p>There&#39;s at least one public data point. <a href=\"https://github.com/antirez/ds4\" rel=\"nofollow ugc noopener\">DwarfStar&#39;s README</a> says it&#39;s developed &quot;with strong assistance from AI coding agents,&quot; with humans leading the ideas, testing and debugging. I don&#39;t know how much of Strata was agent-written. My guess is that the cost of a purpose-built runtime has dropped from months of specialist time to weeks or less, but I don&#39;t have hard numbers.</p>\n<h3 id=\"3-make-tok-s-go-up-is-an-unusually-crisp-objective\">3. &quot;Make tok/s go up&quot; is an unusually crisp objective</h3>\n<p>Most software resists automation because the goal is fuzzy. Inference optimization is closer to machine-checkable:</p>\n<ol><li><strong>Input:</strong> a fixed weights file and token IDs.</li><li><strong>Constraint:</strong> outputs stay within a tolerance of a reference, for example greedy-token agreement or low KL divergence on reference logits.</li><li><strong>Objective:</strong> more tokens per second in less memory.</li></ol>\n<p>You can point an agent loop at that and let it iterate against hardware counters.</p>\n<p>It isn&#39;t fully specified, though. Prefill latency, power, concurrency, context length, sampling behavior and reliability all matter, and a loose tolerance invites the loop to trade away quality it doesn&#39;t measure, such as long-context accuracy or rare tokens. The tighter the parity check, the more you can trust the result. That&#39;s also why the quality check is the most important open item in my Strata write-up.</p>\n<h2 id=\"the-twist-general-engines-become-the-parts-bin\">The twist: general engines become the parts bin</h2>\n<p>It would be easy to frame this as napkin runtimes versus general runtimes. The engines themselves suggest something messier.</p>\n<p>DwarfStar borrows kernels and quant formats from llama.cpp&#39;s GGML. llamAmpere is a llama.cpp fork. The GGUF format, the quant types and many of the kernels these projects start from came out of years of general-purpose work.</p>\n<p>So the more likely future isn&#39;t one replacing the other. General engines keep doing the slow, broad work: formats, quantization research, correct reference kernels for every backend. Napkin runtimes take those parts, bolt them into something overfit for one model and one card, and throw the assembly away when the next model arrives.</p>\n<p>General engines become the parts bin. Napkin runtimes are what power users build from it.</p>\n<h2 id=\"the-console-analogy\">The console analogy</h2>\n<p>Game developers already know this pattern. A console studio targets one exact chip, one memory bus and one cache hierarchy, and that fixed target gives it permission to specialize aggressively. PC developers can&#39;t, because their engine has to run on whatever hardware shows up, so they pay a generality tax in abstraction layers and defensive code paths.</p>\n<p>{% image_cc &quot;./src/static/img/diagrams/overfit-console-vs-pc.png&quot;, &quot;Two-column comparison. PC Model, Generality Tax: a stack of llama.cpp / vLLM / Ollama, Graph schedulers, Multi-OS, CUDA / ROCm / CPU and 50+ model families. Console Model, Direct-to-Metal: Strata / ninfer / Splash pointing straight at Ada SM89 / PCIe 4.0.&quot;, &quot;sketch-draw&quot;, &quot;PC model vs console model of inference runtimes. A generalist runtime pays for graph schedulers, multi-OS and multi-vendor support, and dozens of model families. An overfit runtime is tuned for one target and drops everything else.&quot; %}</p>\n<p>For years nobody wrote console-style engines for a single PC configuration, because no one could justify the engineering cost for an audience of one card. Cheaper engine-writing changes that math. If a runtime for one model on one GPU costs weeks instead of years, your desktop can be treated like a console: a fixed target worth specializing for.</p>\n<h2 id=\"what-survives-the-churn\">What survives the churn</h2>\n<p>If runtimes are disposable, the layers around them have to be stable. Nobody wants a new CLI, UI and SDK every time an engine dies.</p>\n<h3 id=\"the-http-api\">The HTTP API</h3>\n<p>The OpenAI Chat Completions API and the Anthropic Messages API are what clients speak. As long as a napkin runtime exposes one of them, it plugs into Cursor, <a href=\"https://github.com/cline/cline\" rel=\"nofollow ugc noopener\">Cline</a>, Open WebUI, Aider and the rest without anyone noticing what&#39;s underneath.</p>\n<h3 id=\"model-routers\">Model routers</h3>\n<p>Running several one-off engines means juggling processes and ports. That&#39;s what <a href=\"https://github.com/mostlygeek/llama-swap\" rel=\"nofollow ugc noopener\"><code>llama-swap</code></a> and my <a href=\"https://github.com/carteakey/l3ms\" rel=\"nofollow ugc noopener\">L3MS</a> supervisor are for. My current tiers:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model name</th><th>Engine</th><th>Role</th><th>Decode</th></tr></thead><tbody><tr><td><code>qwen38-flash-next-plat</code></td><td>Strata v0.1.30</td><td>Fast default, 128k context</td><td>~53–60 tok/s</td></tr><tr><td><code>qwen38-flash-next-plat-vision</code></td><td>Strata v0.1.30 + CPU vision encoder</td><td>Fast image tasks</td><td>Untested with images; the CPU encoder mostly adds time to first token</td></tr><tr><td><code>qwen38-flash-next</code></td><td>llama.cpp master</td><td>Stable fallback, 96k context</td><td>20.8 tok/s</td></tr><tr><td><code>qwen38-flash-next-vision</code></td><td>llama.cpp master + mmproj</td><td>Image tasks fallback, 16k context</td><td>18.2–18.6 tok/s</td></tr><tr><td><code>qwen38-flash-next-mtp</code></td><td>llama.cpp MTP branch</td><td>Experimental</td><td>25.3–27.1 tok/s</td></tr><tr><td><code>gemma-4-26b-qat-mtp</code></td><td>llama.cpp + MTP assistant</td><td>Smaller MoE, mostly on the GPU</td><td><a href=\"/blog/local-inference/gemma-4-26b-qat-mtp/\">~100 tok/s</a></td></tr></tbody></table></div>\n<p>llama-swap listens on one port. When a client asks for a model, it stops whatever was loaded, starts the right engine with its flags, proxies the request, and shuts it down after 10 idle minutes. The client sees one endpoint; engines swap underneath like cartridges. Trying Strata didn&#39;t mean replacing my workflow, and llama.cpp is still one model name away when the experiment breaks.</p>\n<h3 id=\"hardware-specific-communities-speculative\">Hardware-specific communities (speculative)</h3>\n<p>I&#39;d expect tuning groups to form around exact hardware combinations rather than broad forums. Think a community for a 24 GB card with 64 GB of DDR5, or for one Mac memory tier, trading engine builds, expert routing profiles and placement recipes for their exact setup. This is extrapolation; I haven&#39;t seen it happen yet.</p>\n<h2 id=\"the-cost-trust\">The cost: trust</h2>\n<p>The same properties that make napkin runtimes fast make them risky to run. They&#39;re young, maintained by one person or a small group, change daily, and ship native code that pins memory and talks straight to the GPU driver. Few people review them, and nobody issues security advisories for a project that will be abandoned in six months.</p>\n<p>If this becomes the normal way to run local models, power users take on a supply-chain problem they didn&#39;t have with a well-reviewed general runtime. A source build and a pinned revision make the work inspectable; they don&#39;t make it audited. The minimum I&#39;d suggest: build from source, pin a commit, skim the diff before updating, and don&#39;t run them on machines that hold anything you can&#39;t afford to lose. Hardware communities that share vetted builds could help here, or make it worse.</p>\n<h2 id=\"how-general-engines-could-win-back-the-lead\">How general engines could win back the lead</h2>\n<p>I don&#39;t think general engines go away. Three ways they could close the gap:</p>\n<p><strong>Recipes as plugins.</strong> Alongside a model file, you download a recipe for your hardware: fused kernels, an expert placement mask and a speculation schedule built for, say, an RTX 4070 with DDR5. The general engine becomes a thin, trusted host that loads it. This is the parts-bin idea run in reverse, and it would also ease the trust problem, since the host stays reviewed. It&#39;s a proposed design; I&#39;m not aware of a project that ships this today.</p>\n<p><strong>Agents maintaining the variants upstream.</strong> The same agents that write one-off engines could generate and test hardware-specialized kernel variants inside a general repo, for every GPU generation, without maintainers hand-writing each one. The bottleneck shifts from writing code to reviewing it.</p>\n<p><strong>Compilers finally deliver.</strong> If Mojo/MAX, Triton or MLIR-based stacks get good enough to take a high-level model graph and a hardware description and emit near-optimal code automatically, hand-overfit engines lose their edge overnight. Until then, a person or agent writing for one exact GPU will usually beat compiler output.</p>\n<h2 id=\"don-t-marry-your-inference-engine\">Don&#39;t marry your inference engine</h2>\n<p>For two years, running local models meant one playbook: install llama.cpp or vLLM, pick a context size, offload what fits. That playbook still works, and for most people it&#39;s still the right one.</p>\n<p>But when a runtime written for one model and one GPU doubles your best decode speed and holds it flat across 60,000 tokens, the generality tax gets hard to ignore. So keep your serving layer stable, treat the engine underneath as replaceable, and don&#39;t expect this year&#39;s fastest runtime to support next year&#39;s model.</p>\n<p>Napkin software: cheap to write, very good at the meal in front of you, and fine to throw away when the next course arrives.</p>\n<p>{% callout &quot;note&quot;, &quot;What this argument rests on&quot; %}\nOne model and one machine that I measured myself, with a missing ablation and no quality A/B yet, plus five other projects whose claims I&#39;m taking at face value. The trend is real enough to notice. Whether it lasts is what the April 2027 check is for. I&#39;d like to hear where you think it breaks.\n{% endcallout %}</p>\n<h2 id=\"changelog\">Changelog</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>Date</th><th>Note</th></tr></thead><tbody><tr><td>2026-10-03</td><td>Moved the benchmarks, memory layout and reproduction details to a separate <a href=\"/blog/local-inference/strata-on-an-rtx-4070/\">Strata post</a>. Reframed the headline speedup against the best llama.cpp setup (about 2×), added a checkable prediction, the parts-bin section and a trust section, softened the objective-function claim, trimmed the console analogy, updated the DwarfStar description, and dropped an unverified reference.</td></tr><tr><td>2026-10-01</td><td>Corrected the bandwidth arithmetic and removed the double-counted cache and MTP gain, aligned the hardware specs, linked all six engines, hedged the unsourced claims, moved the IQ3_S and NAS-archive notes to the Flash-Next post, and tightened the prose.</td></tr><tr><td>2026-09-30</td><td>Initial post.</td></tr></tbody></table></div>","headings":[{"level":2,"text":"The surprise","id":"the-surprise"},{"level":2,"text":"The new breed","id":"the-new-breed"},{"level":2,"text":"A prediction you can check","id":"a-prediction-you-can-check"},{"level":2,"text":"Why now: three reasons","id":"why-now-three-reasons"},{"level":3,"text":"1. Generality is a performance tax","id":"1-generality-is-a-performance-tax"},{"level":3,"text":"2. AI coding made engines cheap to write","id":"2-ai-coding-made-engines-cheap-to-write"},{"level":3,"text":"3. \"Make tok/s go up\" is an unusually crisp objective","id":"3-make-tok-s-go-up-is-an-unusually-crisp-objective"},{"level":2,"text":"The twist: general engines become the parts bin","id":"the-twist-general-engines-become-the-parts-bin"},{"level":2,"text":"The console analogy","id":"the-console-analogy"},{"level":2,"text":"What survives the churn","id":"what-survives-the-churn"},{"level":3,"text":"The HTTP API","id":"the-http-api"},{"level":3,"text":"Model routers","id":"model-routers"},{"level":3,"text":"Hardware-specific communities (speculative)","id":"hardware-specific-communities-speculative"},{"level":2,"text":"The cost: trust","id":"the-cost-trust"},{"level":2,"text":"How general engines could win back the lead","id":"how-general-engines-could-win-back-the-lead"},{"level":2,"text":"Don't marry your inference engine","id":"don-t-marry-your-inference-engine"},{"level":2,"text":"Changelog","id":"changelog"}]}}