Topic
Everything filed under Performance, newest first.
RSS · JSON · All topics
Vicent Martí explains why hosting Git at scale is hard, how centralized workflows clash with Git’s distributed design, and what Cursor learned about repository hosting performance and architecture.
23 min · 5,236 words
How did AMD Ryzen get 50% faster in two years?
Daniel Lemire compares AMD Ryzen 7 X3D chips from Zen 3 to Zen 5: Geekbench gains of ~50% came less from clock and more from wider cores, larger caches, bigger ROBs, and 512-bit SIMD.
2 min · 468 words
Persistent Databases in the Browser with DuckDB-Wasm and OPFS
DuckDB explains how DuckDB-Wasm can open a persistent database file in the browser’s Origin Private File System (OPFS), when data reaches disk, and how that changes browser analytics apps that previously relied on Parquet-in-IndexedDB workarounds.
8 min · 1,800 words
Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
PrismML introduces Bonsai 2 27B, a near-lossless compression of a 27B-class multimodal model into roughly a 9× smaller footprint aimed at efficient on-device and local inference.
4 min · 931 words
How Uber Protects Against Retry StormsError ownership so retries know when they help—and when they make outages worse
Uber Engineering explains retry storms in deep service graphs and how context-aware error ownership, claim headers, and middleware stop retries from amplifying a single downstream failure across the stack.
10 min · 2,292 words
Small Programming Tricks Matter
Day to day, I think a surprising amount of engineering productivity comes from small nuggets of knowledge: being aware that a language feature exists; knowing that an unexplained tcp delay is probably related to the TCPNODELAY setting and Nagle’s algorithm; knowing the right git incantation to get out of a pickle; or knowing a trick with sed to rewrite a file. In one sense, this is self-evident: anything you know is going to be made up of smaller pieces of knowledge. Of…
3 min · 772 words
Welcome to the first feature article on our site. We’re going to cover an ongoing problem with x86 emulation that affects every application that we emulate. This comes down to a single over-arching term that has wide-reaching ramifications; Emulating the x86 Total Store Ordering memory model (x86-TSO).
31 min · 7,172 words
Introducing TIN: full-text search for Postgres
PlanetScale announces TIN (Text INdex), a GA full-text search extension for Postgres and Neki with boolean/phrase/span queries, fuzzy and regex matching, BM25 ranking, and transaction-correct updates—built to be fast while staying inside Postgres.
15 min · 3,349 words
Size-Specialized Memory Allocation
Go 1.27 includes faster memory allocation for allocations of 80 bytes or fewer. Allocations can be up to 20-30% faster, making allocation-heavy programs up to 1% faster. The Go runtime improves the performance of those allocations by adding specialized functions that are used to allocate certain sizes. These specialized functions can then make certain assumptions that make them faster and easier to optimize. This blog post will explain how this works and how it makes your programs faster. Heap allocations are created by the runtime’s mallocgc function, which requires the…
7 min · 1,560 words
The query finished… Why is my Fabric SQL database still consuming CUs?
Two minutes of SQL database activity in Microsoft Fabric can mean ~17 minutes of compute billing. Nikola Ilic walks through CU metering with application, development, and troubleshooting examples.
8 min · 1,796 words
Migrating the GitHub Copilot runtime to Rust, using Copilot
Stephen Toub recounts porting GitHub Copilot’s agent runtime from TypeScript/Node to 800k+ lines of production Rust with Copilot agents across 128 incremental PRs, and what the performance and process lessons were.
64 min · 14,699 words
Training a 4B model to produce 81% faster query plans than Postgres
Leis et al. asked this exact question in 2015. Then, they asked it again 10 years later. Despite an enormous body of research spanning a decade since their original exploration, they found that query optimizers continue to leave much to be desired. I was surprised when I first learned about this. A Postgres database should know everything about the stuff that lives in its tables, no? How hard can it be?
41 min · 9,393 words
Subnormal floating-point numbers are expensive… on Intel processors
Daniel Lemire benchmarks IEEE subnormal floating-point performance across Intel Granite/Emerald Rapids, AMD Zen 5, AWS Graviton 5, and Apple M4 Max, finding ~45–50× slower multiplies on Intel while AMD and Arm stay near full speed.
2 min · 466 words
Performance Improvements in .NET 11
Take a tour through hundreds of performance improvements in .NET 11.
164 min · 37,742 words
Better Vector Search for Long Documents: Chunking Inside Manticore Search
An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.
31 min · 7,181 words
The last mile of a long road: faster NumPy in the browser
Notebook.link explains how Emscripten-forge NumPy now links OpenBLAS in WebAssembly, delivering large matmul and linear-algebra speedups for browser scientific computing.
19 min · 4,315 words
You can run git on object storage if you re-make packfiles
Building ObjGit, Tigris explains why Git packfiles fight object storage, how remaking packfiles unlocks workable remote repositories, and what that means for Git servers backed by S3-style buckets.
15 min · 3,408 words
When the fractional part of a float fixes your shader
Bruno Croci debugs a Voronoi Shadertoy stutter that only appeared on a Windows RTX 4070: ANGLE/FXC optimizations dropped `fract` for integer-looking inputs, and swapping a literal to `398.1` or using `x-floor(x)` restored smooth motion.
1 min · 288 words
Principles for fast Tokio applications
A practical guide to building low-latency, high-throughput applications on the Tokio async runtime, written from experience debugging real production systems. The post covers measuring schedule latency, the fairness-versus-batching trade-off, mutex pitfalls, isolating Tokio workers from other threads, and when multiple runtimes make sense.
1 min · 295 wordsagent-written
I made a build profiler to understand Bun's compile times
Lalit Maganti built buildprof, an open-source Linux tool that records every process spawned during a build and renders them on a shared timeline. He used it to investigate the 5x speed difference between Bun's Zig and Rust builds, tracing the bottleneck to Full LTO in the linker and a downloaded WebKit library that amplified it.
1 min · 318 wordsagent-written