Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
August 27 TCRF DDoS Attack Postmortem
The Cutting Room Floor’s postmortem on a sustained August 2026 DDoS: what broke, how mitigation unfolded, and the infrastructure changes made to keep a volunteer game-preservation wiki online.
17 min · 3,930 words
We just shipped support for the ugliest part of HTTP: Vary
Cloudflare Cache Rules now support HTTP Vary on every plan—normalize negotiation headers, pass exact values to origin, or bypass cache when variance is too wild.
12 min · 2,727 words
Understanding NvPCRs in systemd v262
systemd answers TPM PCR scarcity with additional PCR-like registers allocated in the TPM’s NV memory, with an anchoring design that was reworked in v262.
29 min · 6,748 words
Self-hosting LLM models for software development
Kévin Maschtaler on running medium-sized open LLMs on AWS Spot EC2 for day-to-day software work—what stacks, costs, and performance looked like versus a personal Claude subscription.
7 min · 1,719 words
Protobuf, JSON Schema, and OpenAPI
Buf explains how Protobuf schemas can drive JSON Schema and OpenAPI via protoc plugins—extending one source of truth into documentation, validation, and HTTP APIs without maintaining parallel definitions.
6 min · 1,373 words
Tailscale performance updates cut memory use, raise throughput, and speed startup via multi-queue, writev, and netmap caching—how the team measured and shipped the gains.
7 min · 1,640 words
How Google Agent Substrate Works: 250 Agents on 8 Pods
A technical breakdown of Google’s Agent Substrate: how it multiplexes hundreds of stateful agent sessions onto a handful of Kubernetes pods with fast suspend/resume.
11 min · 2,511 words
A file path looks like identity, but it is not: Ryan Galloway explains why treating paths as durable keys breaks pipelines, and what to use instead for content identity.
4 min · 1,016 words
Farid Zakaria on dynamic derivations and nondeterministic build graphs ahead of NixCon 2026—when a build step can roll dice, and what that means for caching and reproducibility.
7 min · 1,636 words
S3 Is the Future, S3 Is the Past
Amazon S3, and its analogues in other clouds, have become the foundation of the modern cloud software architecture. Today, nearly every data-intensive system is being built around S3. However, the hardware assumptions baked into S3’s design – and into all the software architectures that have emerged around it – are rapidly becoming obsolete.
3 min · 742 words
Own the Agent, Rent the Intelligence: Building My Always-On AI Agent Server
James M explains why a Mac mini M6 became his always-on Hermes agent server—routing hard work to cheap cloud models like DeepSeek Flash and Claude Sonnet instead of owning local inference hardware.
24 min · 5,439 words
Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduces AIPerf, the GenAI-Perf successor: a multiprocess LLM inference benchmarker that avoids client bottlenecks at high concurrency, with flexible load shapes, trace replay, and production-scale measurement guidance.
7 min · 1,661 words
Saving another 100TB of RAM with math (and Rust)
Cloudflare explains how math-heavy redesigns and Rust in Pingora cut another ~100TB of RAM across their global network by shrinking hot in-memory structures without sacrificing correctness.
14 min · 3,157 words
Vicent Martí explains why hosting Git at scale is hard, how centralized workflows clash with Git’s distributed design, and what Cursor learned about repository hosting performance and architecture.
23 min · 5,236 words
Telstra outage: The night a network decided the year was 2006
The opposite is actually the case. If we do not have a common understanding of what “now” is, a lot of things we take for granted will stop working.
15 min · 3,430 words
Size-Specialized Memory Allocation
Go 1.27 includes faster memory allocation for allocations of 80 bytes or fewer. Allocations can be up to 20-30% faster, making allocation-heavy programs up to 1% faster. The Go runtime improves the performance of those allocations by adding specialized functions that are used to allocate certain sizes. These specialized functions can then make certain assumptions that make them faster and easier to optimize. This blog post will explain how this works and how it makes your programs faster. Heap allocations are created by the runtime’s mallocgc function, which requires the…
7 min · 1,560 words
The engineering behind the US Strategic Petroleum Reserve
One of the most fascinating things I’ve learned about recently is the engineering behind the US Strategic Petroleum Reserve. Here are the key requirements it’s designed to meet:
4 min · 939 words
Cohere's North Mini Code Megakernel Serving Engine
Today, Cohere presents a serving engine for North Mini Code built around a decode megakernel: BF16 on a single H100, 1.25× - 1.41× faster than vLLM end-to-end. Explore the code behind the serving engine on GitHub. Most LLM serving stacks still treat each forward pass as a sequence of kernels: launch QKV, wait; launch attention, wait; launch the MoE, wait. Each launch is fine on its own. The problem is the waiting in between. At small batch sizes, the GPU spends a surprising fraction of every decode step waiting rather than computing.
24 min · 5,534 words
vLLM x AgentX: Optimizing for Real-World Agentic Serving
**TL;DR:** Agentic workloads are becoming a major source of vLLM traffic. Their multi-turn sessions, long contexts, and extensive prefix reuse demand optimizations across the serving stack. This post walks through vLLM's coordinated approach: KV cache management, parallelism and engine optimizations, and methodologies for prefill/decode disaggregation. Measured on AgentX, SemiAnalysis's public agentic benchmark, vLLM achieves up to 130K total tokens per GPU-second on DeepSeek V4 Pro, and an…
17 min · 3,923 words