Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
OxCaml - Stack Allocations and Locality
The motivation behind OxCaml is to make OCaml a great language for performance engineering, with the eventual goal being to upstream these language extensions to vanilla OCaml (OxCaml](https://oxcaml.org/)). OxCaml maintains backwards compatibility with OCaml, which implies that every OCaml program is a valid OxCaml program. The language extensions range from additions to the type system that rule out data races, to control over allocations that reduces garbage collection pressure, to management
7 min · 1,697 words
PlanetScale Released Text Search and We Have a Lot to Say (Part I)
Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.
14 min · 3,188 words
Context management is an underrated habit
How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work. An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again. Here are the context management techniques we use on the
5 min · 1,257 words
Python 3.15 ships with a new profiler. It is called Tachyon, it lives in the standard library as the profiling.sampling module, and unlike cProfile it is a sampling profiler rather than a tracing one. I now have a set of hands-on workshops for it, which you can find at github.com/GrahamDumpleton/tachyon-workshops or on the workshops page of this site, and they have reached the point where I am happy for other people to do them. That said, they were not written for other people in the first place. They were written so I could learn Tachyon myself, and the reason I wanted...
8 min · 1,868 words
Halfspace: An experimental IDE for solid modeling with distance fields
Matt Keeter’s experimental IDE for solid modeling with distance fields: interactive halfspace CSG, live rendering, and a toolkit for sculpting shapes as signed distance functions.
6 min · 1,448 words
Why we built the fastest robust TTS model
Gradium's latest streaming TTS hits ~50ms time-to-first-audio while improving naturalness and hard cases like phone numbers—freeing latency budget for LLM turns and barge-in in voice agents.
2 min · 431 words
zenkai: The App Launcher I Wrote Because I Wanted Something Fast and Beautiful
Dayvster builds zenkai, a Zig + Qt6 cross-platform app launcher with ~140ms startup (sometimes ~20ms), 65+ themes, Lua plugins, and a sandbox—written as a hobby performance deep dive.
2 min · 571 words
Earendil's experimental Pi Durable harness brings Pi's minimalism to long-running agents: crash-safe tasks, multi-conversation concurrency, pluggable extensions, compaction, durable documents, and multiplayer steering on JS runtimes.
5 min · 1,053 words
India vs West Indies, 3rd ODI: 79 AI models predict the result
We asked 79 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 96% picked India. See every answer, who searched the web, and who dissented.
10 min · 2,281 words
Is a hot dog a sandwich? what 85 AI models think
We asked 85 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 66% picked No. See every answer, who searched the web, and who dissented.
13 min · 2,969 words
Who wins the 2026 F1 drivers' title? 83 AI models predict the result
We asked 83 AI models (GPT, Claude, Gemini, DeepSeek…) the same question. 55% picked Antonelli. See every answer, who searched the web, and who dissented.
10 min · 2,270 words
Cloudflare Containers, rebuilt to scale agent sandboxes
Cloudflare Containers now start about 6x faster, let agents choose each sandbox image and instance type at runtime, and support filesystem snapshots in public beta—controlled from a Durable Object.
14 min · 3,113 words
How to speed up the Rust compiler in September 2026
How to speed up the Rust compiler in September 2026 My last post](https://nnethercote.github.io/2026/07/31/how-to-speed-up-the-rust-compiler-in-july-2026.html) on the Rust compiler’s performance was two months ago and a lot has happened since then. Overall progress The measurements for the period 2026-07-29 to 2026-09-28 can be seen here](https://perf.rust-lang.org/compare.html?start=1a833e16546c2eb012758ddd499964fd8afee29e&stat=wall-time&tab=compile&end=c1070d69382b8d2f2eb65119c738a77d9e324c9e&
5 min · 1,212 words
Garbage Collection: Generational? Incremental? Both!
The third Language Summit talk was brought by Mark Shannon, who is the author of the incremental garbage collector implementation shipped in Python 3.14 that was reverted back to the generational garbage collector from Python 3.13 after reports of “significant memory pressure” in production environments. The original goal of the new incremental garbage collector was to reduce maximum pause times by an order of magnitude for larger heaps.
5 min · 1,173 words
How little does an AI agent need, and how cheap can it get?
Arduino asks what minimum compute an AI agent needs in the physical world—and why cheap Linux+MCU boards change the economics of training and deploying agents at scale.
3 min · 709 words
Edward Kmett's week-old Turbo Haskell Compiler (THC) JITs GHC Core onto Truffle/GraalVM, supports AOT Native Image, polyglot FFI, Loom green threads, and can compile pandoc, happy, alex, and GHC itself.
2 min · 506 words
Browserbase’s Harsehaj Dhami explains Web Bot Auth: cryptographic HTTP message signatures that let AI agents prove identity, while leaving access and reputation decisions to site owners and registries.
5 min · 1,123 words
How our vibe coded website looks like a designer made it
Railcode founder Yakko Majuri walks through a real agent-assisted design process—inspiration spectra, single-HTML forks, color playgrounds, and relentless iteration—showing how non-designers can steer coding agents toward tasteful product sites.
2 min · 397 wordsagent-assisted
What Is a Container, Really? Five Years of GPU Infrastructure
Beam Cloud recounts five years of GPU infrastructure: from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and what “container” actually means in production AI compute.
9 min · 2,132 words