Topic
Everything filed under Performance, newest first.
RSS · JSON · All topics
42x Faster Prompt Lookup Drafting in llama.cpp
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
13 min · 2,954 words
Optimizing x264 Settings and Per-title Ladders
Jan Ozer shows how to tune x264 encodes and per-title bitrate ladders to cut bandwidth while improving quality—practical encoder settings and ladder design for H.264-heavy streaming workflows.
17 min · 3,835 words
Improving site performance by shipping more CSS
GitHub’s Primer team recounts fully migrating github.com off CSS-in-JS to CSS Modules—feature flags, sx-prop cleanup, theming, and the performance wins along the way.
3 min · 723 words
The state of SIMD in Rust in 2026
A lot of progress was made since last year, and I made some of it! After [last year's survey](https://shnatsel.medium.com/the state of simd in rust in 2025 32c263e5f53d) I started contributing to the SIMD library that seemed the most promising. One thing led to another, and now I'm a maintainer of Fearless SIMD. To avoid a conflict of interest, I invited authors of other libraries ( std::simd , wide , pulp , macerator ) to review and provide feedback on a draft of this article.
25 min · 5,668 words
Inline vs. Separate Tables for Vectors in Postgres: Measuring the Join Overhead
Separate Tables for Vectors in Postgres: Measuring the Join Overhead Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost of decoupling your vectors.
10 min · 2,296 words
When to choose x86-64 vs aarch64
Not all cloud vCPUs are created equal. When you create a PlanetScale Postgres or Neki database, you have to choose between `aarch64` (ARM) and `x86-64`. Two clusters on different architectures can have the same vCPU count and RAM, yet perform very differently. It's worth understanding the implications, since you can't easily switch the CPU architecture on your cluster later. x86 grew up on the desktop, prioritizing backward compatibility and performance. It began at Intel in 1978 and IBM’s…
6 min · 1,326 words
Mac Mini M6: Retro PC Emulation with 86Box (600MHz PII?!)
Benchmarking cycle-accurate 86Box on an M6 Mac Mini: why one core matters for retro PC emulation, a stable 600MHz Pentium II clock, and how timing accuracy compares to original hardware.
7 min · 1,678 words
Postgres SELECT DISTINCT Does Not Scale
DBOS explains why Postgres SELECT DISTINCT can get surprisingly expensive as datasets grow, what the planner is doing, and how they mitigated it in practice.
5 min · 1,068 words
Packing Binary Is Fun, Actually
Pranav Desai’s hands-on tour of binary packing: why packing bits can be fun, the techniques that matter, and practical patterns for packing denser structures without losing your mind.
22 min · 5,056 words
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Charles Frye and Shreya on the Modal blog: combining a query planner with an inference engine to push AI-SQL queries past a billion tokens per minute on one GPU—why left-deep joins help KV cache, and how they beat naive vLLM-style serving.
18 min · 4,181 words
Introducing WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL
Today, we’re announcing WalShadow, an open-source engine that replicates Postgres data to ClickHouse directly from physical WAL. In our benchmarks, transactions committed in Postgres became visible in ClickHouse in around 200 ms, while WalShadow sustained 289K rows/sec, effectively keeping pace with the source Postgres instance. Unlike traditional CDC based systems, WalShadow doesn’t use Postgres logical replication. It consumes the same physical WAL stream used by Postgres replicas, decodes…
5 min · 1,059 words
Platform-independent SIMD in Go
Go 1.26 and 1.27 include experimental APIs for Single Instruction Multiple Data SIMD operations. SIMD is a native feature of many modern CPUs that allows software to perform uniform operations across vectors of data very quickly, such as adding 8 pairs of float64 values in a single instruction. It can significantly speed up many computationally-intensive tasks, ranging from cryptography to data processing to AI. In fact, Go’s Green Tea garbage collector/blog/greenteagc even makes use of SIMD to accelerate scanning memory for live objects.
15 min · 3,399 words
Topcoat is pushing the boundary of server applications with Rust
Two monthshttps://tokio.rs/blog/2026-07-22-announcing-topcoat ago, we Julienhttps://github.com/pikaju and Ihttps://github.com/carllerche announced Topcoathttps://github.com/tokio-rs/topcoat, a batteries-included full-stack Rust framework. It includes views, components, mailers, an ORM Toastyhttps://github.com/tokio-rs/toasty, and more. Topcoat aims to make building web apps with Rust as productive as any other language. We have been hard at work shipping features, so it is a good time to talk about what is new.
9 min · 1,958 words
Ten years of tmux, and the 1,495 lines of zsh it cost me
Yogesh Lonkar measured a tmux status bar burning about 15% of a CPU core, then rewrote a decade of forking shell scripts into a leaner setup—and documents what 1,495 lines of zsh had been doing the whole time.
13 min · 2,974 words
How to serve trillions of tokens for trillion-parameter coding agents
Modal explains how it serves coding-agent inference at extreme scale—performance and efficiency techniques for trillion-parameter models generating trillions of tokens, written for teams facing the same workload.
30 min · 6,972 words
SDF vs. MSDF vs. Slug: GPU Text Rendering
Chris Hanson compares SDF, MSDF, Slug, texture atlases, and Rive for GPU text: how signed-distance fields trade sharpness, memory, and shader cost when you draw outlines yourself.
15 min · 3,398 words
ReBarUEFI: Resizable BAR for almost any UEFI systemA DXE driver and tooling to enable Resizable BAR on motherboards that never exposed the option.
Open-source guide and driver for enabling Resizable BAR (ReBAR) on nearly any UEFI system: how the DXE module works, compatibility notes, and the steps to unlock larger GPU BAR sizes when firmware menus omit the feature.
4 min · 880 words
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
6 min · 1,375 words
700 MB/s of Kafka throughput, on Postgres
Profiling Kafgres (Kafka-compatible broker inside Postgres) from ~113 MB/s to ~700 MB/s via cached SPI plans, relaxed commits, separate topic disks, and WaitEventSet socket readiness.
5 min · 1,062 words
Mercury 2.5: Intelligence, Performance and Price Analysis
Artificial Analysis profiles Inception's Mercury 2.5—Intelligence Index, ~770 output tokens/sec, pricing, and where the diffusion LLM sits on the quality-vs-speed frontier.
12 min · 2,677 words