Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
OxCaml - Stack Allocations and Locality
The motivation behind OxCaml is to make OCaml a great language for performance engineering, with the eventual goal being to upstream these language extensions to vanilla OCaml (OxCaml](https://oxcaml.org/)). OxCaml maintains backwards compatibility with OCaml, which implies that every OCaml program is a valid OxCaml program. The language extensions range from additions to the type system that rule out data races, to control over allocations that reduces garbage collection pressure, to management
7 min · 1,697 words
PlanetScale Released Text Search and We Have a Lot to Say (Part I)
Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.
14 min · 3,188 words
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
zenkai: The App Launcher I Wrote Because I Wanted Something Fast and Beautiful
Dayvster builds zenkai, a Zig + Qt6 cross-platform app launcher with ~140ms startup (sometimes ~20ms), 65+ themes, Lua plugins, and a sandbox—written as a hobby performance deep dive.
2 min · 571 words
How to speed up the Rust compiler in September 2026
How to speed up the Rust compiler in September 2026 My last post](https://nnethercote.github.io/2026/07/31/how-to-speed-up-the-rust-compiler-in-july-2026.html) on the Rust compiler’s performance was two months ago and a lot has happened since then. Overall progress The measurements for the period 2026-07-29 to 2026-09-28 can be seen here](https://perf.rust-lang.org/compare.html?start=1a833e16546c2eb012758ddd499964fd8afee29e&stat=wall-time&tab=compile&end=c1070d69382b8d2f2eb65119c738a77d9e324c9e&
5 min · 1,212 words
Garbage Collection: Generational? Incremental? Both!
The third Language Summit talk was brought by Mark Shannon, who is the author of the incremental garbage collector implementation shipped in Python 3.14 that was reverted back to the generational garbage collector from Python 3.13 after reports of “significant memory pressure” in production environments. The original goal of the new incremental garbage collector was to reduce maximum pause times by an order of magnitude for larger heaps.
5 min · 1,173 words
Edward Kmett's week-old Turbo Haskell Compiler (THC) JITs GHC Core onto Truffle/GraalVM, supports AOT Native Image, polyglot FFI, Loom green threads, and can compile pandoc, happy, alex, and GHC itself.
2 min · 506 words
What Is a Container, Really? Five Years of GPU Infrastructure
Beam Cloud recounts five years of GPU infrastructure: from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and what “container” actually means in production AI compute.
9 min · 2,132 words
Deser: Rethinking Rust Serialization
Armin Ronacher revisits Deser, an experimental Rust serialization library that inverts Serde’s visitor recursion into heap-backed sinks/emitters—trading some performance for lossless buffering, composable adapters, XML namespaces, and no stack overflow on deep nests.
2 min · 492 words
5x faster Edge Functions: How we replaced v8 isolates with Firecracker MicroVMs
About a billion Edge Functions run on Netlify every day — Sunweb personalizing pages, LotoQuébec routing traffic on a cookie check, and hundreds of thousands of other sites doing everything from personalization to routing to auth. All of it runs on a full JavaScript runtime that scales with our customers’ traffic.
8 min · 1,844 words
Why refactoring made our biggest file biggerCode entropy, measured, and the ratchet that stops it
Image Horse's AppShell grew from 3,250 to 3,806 lines during a month of intentional extraction. Chris Lane Jones explains measured line-count ratchets, why warnings failed, and why entropy is whack-a-mole you don't win—but can make visible.
5 min · 1,168 words
Eddie Aftandilian ships SafeRE 1.0, a linear-time Java regex library built with agents: differential testing vs the JDK, ReDoS resistance by construction, and performance that now beats JDK and RE2/J on Rebar workloads.
5 min · 1,062 words
GrapheneOS – When an app is slow
A GrapheneOS user digs into why OsmAnd maps feel slower on a Pixel 8 than on stock Android, and how that search led to CoMaps and broader performance trade-offs on hardened phones.
1 min · 245 words
Accelerated Out of Core Shuffling
Benjamin Zaitlen explains RapidsMPF’s reusable out-of-core shuffler for distributed analytics—how spilling turns shuffle OOM headaches into a budgetable resource, and what it takes to push shuffle bandwidth toward terabytes per second.
15 min · 3,466 words
42x Faster Prompt Lookup Drafting in llama.cpp
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
13 min · 2,954 words
Improving site performance by shipping more CSS
GitHub’s Primer team recounts fully migrating github.com off CSS-in-JS to CSS Modules—feature flags, sx-prop cleanup, theming, and the performance wins along the way.
3 min · 723 words
The state of SIMD in Rust in 2026
A lot of progress was made since last year, and I made some of it! After [last year's survey](https://shnatsel.medium.com/the state of simd in rust in 2025 32c263e5f53d) I started contributing to the SIMD library that seemed the most promising. One thing led to another, and now I'm a maintainer of Fearless SIMD. To avoid a conflict of interest, I invited authors of other libraries ( std::simd , wide , pulp , macerator ) to review and provide feedback on a draft of this article.
25 min · 5,668 words
Inline vs. Separate Tables for Vectors in Postgres: Measuring the Join Overhead
Separate Tables for Vectors in Postgres: Measuring the Join Overhead Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost of decoupling your vectors.
10 min · 2,296 words
When to choose x86-64 vs aarch64
Not all cloud vCPUs are created equal. When you create a PlanetScale Postgres or Neki database, you have to choose between `aarch64` (ARM) and `x86-64`. Two clusters on different architectures can have the same vCPU count and RAM, yet perform very differently. It's worth understanding the implications, since you can't easily switch the CPU architecture on your cluster later. x86 grew up on the desktop, prioritizing backward compatibility and performance. It began at Intel in 1978 and IBM’s…
6 min · 1,326 words