Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
Mac Mini M6: Retro PC Emulation with 86Box (600MHz PII?!)
Benchmarking cycle-accurate 86Box on an M6 Mac Mini: why one core matters for retro PC emulation, a stable 600MHz Pentium II clock, and how timing accuracy compares to original hardware.
7 min · 1,678 words
Postgres SELECT DISTINCT Does Not Scale
DBOS explains why Postgres SELECT DISTINCT can get surprisingly expensive as datasets grow, what the planner is doing, and how they mitigated it in practice.
5 min · 1,068 words
Packing Binary Is Fun, Actually
Pranav Desai’s hands-on tour of binary packing: why packing bits can be fun, the techniques that matter, and practical patterns for packing denser structures without losing your mind.
22 min · 5,056 words
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Charles Frye and Shreya on the Modal blog: combining a query planner with an inference engine to push AI-SQL queries past a billion tokens per minute on one GPU—why left-deep joins help KV cache, and how they beat naive vLLM-style serving.
18 min · 4,181 words
Platform-independent SIMD in Go
Go 1.26 and 1.27 include experimental APIs for Single Instruction Multiple Data SIMD operations. SIMD is a native feature of many modern CPUs that allows software to perform uniform operations across vectors of data very quickly, such as adding 8 pairs of float64 values in a single instruction. It can significantly speed up many computationally-intensive tasks, ranging from cryptography to data processing to AI. In fact, Go’s Green Tea garbage collector/blog/greenteagc even makes use of SIMD to accelerate scanning memory for live objects.
15 min · 3,399 words
Topcoat is pushing the boundary of server applications with Rust
Two monthshttps://tokio.rs/blog/2026-07-22-announcing-topcoat ago, we Julienhttps://github.com/pikaju and Ihttps://github.com/carllerche announced Topcoathttps://github.com/tokio-rs/topcoat, a batteries-included full-stack Rust framework. It includes views, components, mailers, an ORM Toastyhttps://github.com/tokio-rs/toasty, and more. Topcoat aims to make building web apps with Rust as productive as any other language. We have been hard at work shipping features, so it is a good time to talk about what is new.
9 min · 1,958 words
Ten years of tmux, and the 1,495 lines of zsh it cost me
Yogesh Lonkar measured a tmux status bar burning about 15% of a CPU core, then rewrote a decade of forking shell scripts into a leaner setup—and documents what 1,495 lines of zsh had been doing the whole time.
13 min · 2,974 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
700 MB/s of Kafka throughput, on Postgres
Profiling Kafgres (Kafka-compatible broker inside Postgres) from ~113 MB/s to ~700 MB/s via cached SPI plans, relaxed commits, separate topic disks, and WaitEventSet socket readiness.
5 min · 1,062 words
We just shipped support for the ugliest part of HTTP: Vary
Cloudflare Cache Rules now support HTTP Vary on every plan—normalize negotiation headers, pass exact values to origin, or bypass cache when variance is too wild.
12 min · 2,727 words
Making the MiniMax H3 Video VAE 2x Faster
The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4 2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129 frame encode and decode round trip drops from 24.3 to 12.7 seconds. What changed, the technical details A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes.
2 min · 454 words
Tailscale performance updates cut memory use, raise throughput, and speed startup via multi-queue, writev, and netmap caching—how the team measured and shipped the gains.
7 min · 1,640 words
How we made claude.ai 3x faster in two weeks
Anthropic’s performance sprint cut claude.ai and desktop p75 time-to-typeable from 3.1s to 0.55s: Claude Tag measured journeys, built benchmarks, and shipped thousands of guarded changes in Slack-driven loops.
18 min · 4,175 words
EXPLAIN (ANALYZE, IO) in PostgreSQL 19
Franck Pachot walks through PostgreSQL 19's new EXPLAIN IO stats—prefetch depth, request size, concurrency, and waits—using Little's Law to interpret async read streams.
4 min · 1,031 words
Reducing Image Generation cost with AMD and the Luminal Compiler
Luminal engineers show Flux.2 Klein 9B image generation costs cut by up to 47% on AMD MI300X versus an Nvidia H200, using the Luminal compiler.
15 min · 3,489 words
AI coding has made CI a bottleneck, so we reworked ours to keep up
Linear's Mufeez Amjad explains how agent-accelerated shipping made CI the bottleneck, and how they cut PR wait time and runner cost while test suites nearly quadrupled.
8 min · 1,907 words
Rebuilding Nym’s agent around Jev
How Nym rebuilt its agent stack around TypeSafe’s Jev for guardrails, browser actions, and tool selection—with benchmarks and a shopping demo.
10 min · 2,186 words
Solving for faster SHA-1 collision detection
tl;dr: I discovered collision-detecting SHA-1 is slow and decided to build my own. sha1dc is a rewrite of SHA-1 with collision detection, whose code generator uses a solver to fit collision tests into SIMD lanes. It runs at 68–81% of plain SHA-1's speed where the existing crate runs at 28–29%, and can make git pack verification twice as fast.
10 min · 2,227 words
S3 Is the Future, S3 Is the Past
Amazon S3, and its analogues in other clouds, have become the foundation of the modern cloud software architecture. Today, nearly every data-intensive system is being built around S3. However, the hardware assumptions baked into S3’s design – and into all the software architectures that have emerged around it – are rapidly becoming obsolete.
3 min · 742 words
Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduces AIPerf, the GenAI-Perf successor: a multiprocess LLM inference benchmarker that avoids client bottlenecks at high concurrency, with flexible load shapes, trace replay, and production-scale measurement guidance.
7 min · 1,661 words