Topic
Everything filed under Databases, newest first.
RSS · JSON · All topics
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
6 min · 1,334 words
Announcing Cloudflare K2: serverless event streams
Cloudflare K2 is a serverless event streaming service built directly on top of R2 object storage for high-scale data movement and long-term retention. By decoupling producers and consumers at the edge, K2 enables durable, ordered log streams without the operational overhead of traditional broker clusters.
7 min · 1,546 words
PlanetScale Released Text Search and We Have a Lot to Say (Part I)
Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.
14 min · 3,188 words
Postgres AT TIME ZONE 'UTC' does NOT do what you think it does
A practical Postgres footgun: AT TIME ZONE 'UTC' converts timestamptz to timestamp without time zone—so month math and equality break unless you apply AT TIME ZONE twice.
2 min · 538 words
Inline vs. Separate Tables for Vectors in Postgres: Measuring the Join Overhead
Separate Tables for Vectors in Postgres: Measuring the Join Overhead Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost of decoupling your vectors.
10 min · 2,296 words
When to choose x86-64 vs aarch64
Not all cloud vCPUs are created equal. When you create a PlanetScale Postgres or Neki database, you have to choose between `aarch64` (ARM) and `x86-64`. Two clusters on different architectures can have the same vCPU count and RAM, yet perform very differently. It's worth understanding the implications, since you can't easily switch the CPU architecture on your cluster later. x86 grew up on the desktop, prioritizing backward compatibility and performance. It began at Intel in 1978 and IBM’s…
6 min · 1,326 words
Postgres SELECT DISTINCT Does Not Scale
DBOS explains why Postgres SELECT DISTINCT can get surprisingly expensive as datasets grow, what the planner is doing, and how they mitigated it in practice.
5 min · 1,068 words
Introducing WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL
Today, we’re announcing WalShadow, an open-source engine that replicates Postgres data to ClickHouse directly from physical WAL. In our benchmarks, transactions committed in Postgres became visible in ClickHouse in around 200 ms, while WalShadow sustained 289K rows/sec, effectively keeping pace with the source Postgres instance. Unlike traditional CDC based systems, WalShadow doesn’t use Postgres logical replication. It consumes the same physical WAL stream used by Postgres replicas, decodes…
5 min · 1,059 words
700 MB/s of Kafka throughput, on Postgres
Profiling Kafgres (Kafka-compatible broker inside Postgres) from ~113 MB/s to ~700 MB/s via cached SPI plans, relaxed commits, separate topic disks, and WaitEventSet socket readiness.
5 min · 1,062 words
Why didn't anybody tell me about hash slots
A delivery-matching engineer discovers Redis Cluster hash slots the hard way, and walks through how slot-aware keys change caching, sharding, and multi-key operations in production.
8 min · 1,937 words
EXPLAIN (ANALYZE, IO) in PostgreSQL 19
Franck Pachot walks through PostgreSQL 19's new EXPLAIN IO stats—prefetch depth, request size, concurrency, and waits—using Little's Law to interpret async read streams.
4 min · 1,031 words
PgDog’s Lev Kokotov on why engineers fix open problems they care about, why ownership beats closed-source support tickets, and how free-as-in-freedom Postgres sharding funds itself with enterprise SLAs.
2 min · 387 words
Apache Cassandra® 6 Accord transactions: What you need to know
There have always been architectural trade-offs when considering a distributed database like Apache Cassandra versus a relational database. Cassandra excels at linear horizontal scalability, multi-region replication, and fault-tolerant uptime that relational systems couldn’t match.
7 min · 1,612 words
S3 Is the Future, S3 Is the Past
Amazon S3, and its analogues in other clouds, have become the foundation of the modern cloud software architecture. Today, nearly every data-intensive system is being built around S3. However, the hardware assumptions baked into S3’s design – and into all the software architectures that have emerged around it – are rapidly becoming obsolete.
3 min · 742 words
Persistent Databases in the Browser with DuckDB-Wasm and OPFS
DuckDB explains how DuckDB-Wasm can open a persistent database file in the browser’s Origin Private File System (OPFS), when data reaches disk, and how that changes browser analytics apps that previously relied on Parquet-in-IndexedDB workarounds.
8 min · 1,800 words
Introducing TIN: full-text search for Postgres
PlanetScale announces TIN (Text INdex), a GA full-text search extension for Postgres and Neki with boolean/phrase/span queries, fuzzy and regex matching, BM25 ranking, and transaction-correct updates—built to be fast while staying inside Postgres.
15 min · 3,349 words
The query finished… Why is my Fabric SQL database still consuming CUs?
Two minutes of SQL database activity in Microsoft Fabric can mean ~17 minutes of compute billing. Nikola Ilic walks through CU metering with application, development, and troubleshooting examples.
8 min · 1,796 words
Training a 4B model to produce 81% faster query plans than Postgres
Leis et al. asked this exact question in 2015. Then, they asked it again 10 years later. Despite an enormous body of research spanning a decade since their original exploration, they found that query optimizers continue to leave much to be desired. I was surprised when I first learned about this. A Postgres database should know everything about the stuff that lives in its tables, no? How hard can it be?
41 min · 9,393 words
Appwrite Init 2026 recap: Everything we shipped
Appwrite wraps Init 2026 with a full recap of five days of launches: Appwrite 2.0, native PostgreSQL and MySQL, VectorsDB, S3-compatible storage, Firewall, OAuth2 server, and custom Domains.
7 min · 1,505 words
Better Vector Search for Long Documents: Chunking Inside Manticore Search
An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.
31 min · 7,181 words