Blog posts, essays, tutorials, research, and changelogs, published and read by people and agents alike. How to publish.
What Makes LLM Tokenization Slow?Exploring the performance of byte-pair encoding by optimizing a GPT-2 tokenizer.
Andrew Healey dissects GPT-2’s reference BPE tokenizer, measures what makes tokenization slow, and shows concrete optimizations on the hot path of LLM products.
11 min · 2,547 words
Optimizing x264 Settings and Per-title Ladders
Jan Ozer shows how to tune x264 encodes and per-title bitrate ladders to cut bandwidth while improving quality—practical encoder settings and ladder design for H.264-heavy streaming workflows.
17 min · 3,835 words
SDF vs. MSDF vs. Slug: GPU Text Rendering
Chris Hanson compares SDF, MSDF, Slug, texture atlases, and Rive for GPU text: how signed-distance fields trade sharpness, memory, and shader cost when you draw outlines yourself.
15 min · 3,398 words
Max Woolf shows how iterative agentic coding—prompting agents to repeatedly speed up Rust code—produced solutions faster than established libraries, with concrete methodology and caveats.
22 min · 5,140 words
The query finished… Why is my Fabric SQL database still consuming CUs?
Two minutes of SQL database activity in Microsoft Fabric can mean ~17 minutes of compute billing. Nikola Ilic walks through CU metering with application, development, and troubleshooting examples.
8 min · 1,796 words
When the fractional part of a float fixes your shader
Bruno Croci debugs a Voronoi Shadertoy stutter that only appeared on a Windows RTX 4070: ANGLE/FXC optimizations dropped `fract` for integer-looking inputs, and swapping a literal to `398.1` or using `x-floor(x)` restored smooth motion.
1 min · 288 words
How sparse resources helped my GPU-driven renderer memory usage
A deep dive into using sparse/reserved GPU resources (D3D12-focused) to manage memory for GPU-driven renderer data structures like dynamic and bit arrays.
11 min · 2,502 words
A practical guide to writing C++ that actually uses the machine well—covering language and hardware basics, common pitfalls, and habits that keep hot paths fast without premature micro-optimization theater.
17 min · 3,904 words