Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
Python 3.15 ships with a new profiler. It is called Tachyon, it lives in the standard library as the profiling.sampling module, and unlike cProfile it is a sampling profiler rather than a tracing one. I now have a set of hands-on workshops for it, which you can find at github.com/GrahamDumpleton/tachyon-workshops or on the workshops page of this site, and they have reached the point where I am happy for other people to do them. That said, they were not written for other people in the first place. They were written so I could learn Tachyon myself, and the reason I wanted...
8 min · 1,868 words
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
Earendil's experimental Pi Durable harness brings Pi's minimalism to long-running agents: crash-safe tasks, multi-conversation concurrency, pluggable extensions, compaction, durable documents, and multiplayer steering on JS runtimes.
5 min · 1,053 words
Cloudflare Containers, rebuilt to scale agent sandboxes
Cloudflare Containers now start about 6x faster, let agents choose each sandbox image and instance type at runtime, and support filesystem snapshots in public beta—controlled from a Durable Object.
14 min · 3,113 words
Browserbase’s Harsehaj Dhami explains Web Bot Auth: cryptographic HTTP message signatures that let AI agents prove identity, while leaving access and reputation decisions to site owners and registries.
5 min · 1,123 words
What Is a Container, Really? Five Years of GPU Infrastructure
Beam Cloud recounts five years of GPU infrastructure: from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and what “container” actually means in production AI compute.
9 min · 2,132 words
5x faster Edge Functions: How we replaced v8 isolates with Firecracker MicroVMs
About a billion Edge Functions run on Netlify every day — Sunweb personalizing pages, LotoQuébec routing traffic on a cookie check, and hundreds of thousands of other sites doing everything from personalization to routing to auth. All of it runs on a full JavaScript runtime that scales with our customers’ traffic.
8 min · 1,844 words
Backblaze Drive Stats for Q2 2026
Backblaze’s Q2 2026 Drive Stats: 354,415 drives at 1.73% quarterly AFR, lifetime AFR 1.41%, plus a clear CMR vs SMR primer as 20TB+ drives become a larger share of the fleet.
11 min · 2,548 words
Why I Stopped Defaulting to Next.js and Vercel
How AI coding agents made it practical for me to build and own a different stack with TanStack Start and Cloudflare. The first person who introduced me to Next.js was my friend Haythem Lazaar. We were at university, building Collo , a project management tool for remote teams.
8 min · 1,755 words
Add Runtime Controls to AI Agents with NVIDIA OpenShell
NVIDIA’s technical write-up on OpenShell: an open secure runtime that sandboxes AI agents, enforces tool/file/network policy at runtime, and pairs with hardware monitoring for containment.
7 min · 1,593 words
Accelerated Out of Core Shuffling
Benjamin Zaitlen explains RapidsMPF’s reusable out-of-core shuffler for distributed analytics—how spilling turns shuffle OOM headaches into a budgetable resource, and what it takes to push shuffle bandwidth toward terabytes per second.
15 min · 3,466 words
Introducing Casita: A content-addressed store for source code and build artifacts
Domen Kožar introduces Casita, a content-addressed object store for source code and build artifacts—an argument for rethinking Nix-style packaging from the package manager up.
7 min · 1,723 words
Deploying Guix images on Linode
David Thompson documents migrating to Linode (Akamai Cloud) and deploying Guix system images—image building, boot configuration, and practical notes from moving off DigitalOcean.
4 min · 1,000 words
John D. Cook explains dawn-dusk sun-synchronous orbits: why satellites there stay in perpetual sunlight, and what that means for powering orbital compute servers.
1 min · 257 words
How I Could've Accessed 17 Trillion Microsoft Records
A security write-up estimating ~17.3 trillion stored rows across Microsoft datasets and showing how misconfigured access paths could have exposed enormous volumes of tenant data—plus responsible disclosure notes.
9 min · 2,086 words
When to choose x86-64 vs aarch64
Not all cloud vCPUs are created equal. When you create a PlanetScale Postgres or Neki database, you have to choose between `aarch64` (ARM) and `x86-64`. Two clusters on different architectures can have the same vCPU count and RAM, yet perform very differently. It's worth understanding the implications, since you can't easily switch the CPU architecture on your cluster later. x86 grew up on the desktop, prioritizing backward compatibility and performance. It began at Intel in 1978 and IBM’s…
6 min · 1,326 words
Every package is already installed
Farid Zakaria introduces omnibin: a FUSE filesystem that puts every binary Nixpkgs ever shipped on your $PATH—nothing installed upfront, 0 bytes on disk until something is actually run.
4 min · 1,001 words
How Cloudflare addressed a cross-tenant data exposure vulnerability in Containers
Cloudflare details how a Containers/Sandboxes cross-tenant bug let residual dm-thin disk blocks leak between customers, how Oren Yomtov reported it, and the fleet-wide remediation completed by 19 Sep 2026.
7 min · 1,583 words
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Charles Frye and Shreya on the Modal blog: combining a query planner with an inference engine to push AI-SQL queries past a billion tokens per minute on one GPU—why left-deep joins help KV cache, and how they beat naive vLLM-style serving.
18 min · 4,181 words