Topic
Everything filed under Data Engineering, newest first.
RSS · JSON · All topics
Accelerated Out of Core Shuffling
Benjamin Zaitlen explains RapidsMPF’s reusable out-of-core shuffler for distributed analytics—how spilling turns shuffle OOM headaches into a budgetable resource, and what it takes to push shuffle bandwidth toward terabytes per second.
15 min · 3,466 words
What happens when you analyze college football like the CIA?
What happens when you analyze college football like the CIA? A couple weekends ago, the Illinois football team lost to Duke at home, 31–27.
11 min · 2,518 words
Inline vs. Separate Tables for Vectors in Postgres: Measuring the Join Overhead
Separate Tables for Vectors in Postgres: Measuring the Join Overhead Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost of decoupling your vectors.
10 min · 2,296 words
Ghost Jobs Report, September 2026
Unlisted finds 28.3% of 607,050 career-site job postings have been open over 90 days—age by category, country, ATS, and the employers with the most stale listings.
7 min · 1,617 words
Writing Parquet Files Using Haskell
A practical walkthrough of generating Apache Parquet from Haskell: schema encoding, column chunks, and the tradeoffs of building interoperable analytical data files outside the JVM ecosystem.
11 min · 2,537 words
Aleksandar Filipovski, 2026-09-16 See also: John Salvatier’s excellent blog, Reality has a surprising amount of detail I read a comment somewhere that stuck with me, that went something like this:
8 min · 1,784 words
The query finished… Why is my Fabric SQL database still consuming CUs?
Two minutes of SQL database activity in Microsoft Fabric can mean ~17 minutes of compute billing. Nikola Ilic walks through CU metering with application, development, and troubleshooting examples.
8 min · 1,796 words
Beyond the model: Engineering AI infra with scientific judgementHow Airbnb's agent harness encodes scientific methodology for unstructured data exploration.
Ask a coding agent to analyze 100,000 customer support conversations and within minutes you’ll have a polished taxonomy, precise prevalence numbers, and an executive-ready summary. What you can’t see is the investigation that produced them: the methods it chose, the evidence it weighed, how much to trust it, or whether a second request would agree. All that reaches you is the polish. The model is undeniably intelligent, but intelligence without methodology is not science.
5 min · 1,106 words
LLM Classification Is Feature Engineering
Taylor Pospisil argues LLMs work better as feature generators than as end-to-end classifiers, covering calibration, thresholding, cost, and how to treat model outputs as engineered features.
13 min · 3,039 words
Eddie argues that Pandas forces data practitioners to adopt distributed compute infrastructure long before their data sizes justify it, and that DuckDB and Polars can fill the gap for workloads up to around 100 GB on a single machine. Benchmark results show Polars and DuckDB completing a one-billion-row task in under a minute, while Pandas takes over twelve.
1 min · 268 wordsagent-written
Planet Labs' Open Satellite Feed
Mark Litwintschik surveys Planet Labs' satellite fleet and walks through practical use of the company's freely available Disaster Data feed, which publishes pre- and post-event imagery from natural disasters as DuckDB-friendly Parquet files hosted on Cloudflare.
1 min · 260 wordsagent-written
Introducing chdb Postgres extension: High-performance imports from cloud storage
ClickHouse announces the chdb Postgres extension: fast imports and exports across cloud storage and formats, powered by the embedded ClickHouse engine and usable via COPY-style workflows.
7 min · 1,599 words
Polars is releasing its first 2.0 release candidate, with the major change being that all LazyFrame queries now default to the streaming engine, delivering substantial memory and performance improvements for most users. The version bump is driven by breaking changes to defaults rather than new features, and a migration guide is provided.
1 min · 269 wordsagent-written
No Easy Fix for Bogus Respondents in Online Opt-In Polls
Pew Research Center tests trap questions, CloudResearch Sentry, and voter-file matching on 11,114 opt-in respondents: bogus cases still distort quality, and voter-file matching can raise error by discarding valid people.
16 min · 3,670 words
I've operated petabyte-scale ClickHouse clusters for 5 yearsWhat I learned operating ClickHouse at scale: the wins, the failures, and the lessons that only come from production experience.
Tinybird co-founder Javi Santana shares five years of operational experience running petabyte-scale ClickHouse clusters, covering architecture decisions, the challenges of storage-compute separation, zero-copy replication trade-offs, and how the upgrade process evolved from a three-hour ordeal into a CI/CD-integrated routine.
1 min · 280 wordsagent-written
Alternatives to MinIO for single-node local S3
After MinIO's parent company abandoned the project in late 2025, developer Robin Moffatt evaluated six Docker-first, open-source S3-compatible replacements for use in local demos and data pipeline testing. S3Proxy and SeaweedFS come out as the easiest drop-in alternatives, while Garage and Apache Ozone are judged too complex for single-node lightweight use.
1 min · 265 wordsagent-written
How we built it: Real-time analytics for Stripe BillingOriginal article link with an AI-written directory summary
Reed Trevelyan explains the shift from delayed batch processing to fresher analytics for Stripe Billing. The article describes the architecture and processing changes required to reflect subscription activity while preserving historical accuracy.
1 min · 79 wordsagent-written
How we built it: Jurisdiction resolution for Stripe TaxOriginal article link with an AI-written directory summary
Erich Rentz and Danko Komlen explain how Stripe resolves taxing jurisdictions for transactions. The engineering account explores geographic boundaries, offline preparation, and fast online lookups in a complex and changing ruleset.
1 min · 80 wordsagent-written
Ledger: Stripe’s system for tracking and validating money movementOriginal article link with an AI-written directory summary
Ilya Ganelin describes Ledger, Stripe's internal record of financial movements. The account explains how immutable records and data-quality checks help reconcile expected payment behavior with imperfect external systems.
1 min · 75 wordsagent-written