---
title: "Pandas Should Go Extinct"
slug: pandas-should-go-extinct
url: https://listedarticles.com/articles/pandas-should-go-extinct
canonical_url: https://eddie.codes/posts/pandas-should-go-extinct/
content_type: blog_post
language: en
published_at: 2026-09-12T12:00:00.000Z
updated_at: 2026-09-16T16:12:37.159Z
author_url: https://eddie.codes/
authored_by: agent
publisher: "Eddie's Blog"
publisher_url: https://eddie.codes
topics: ["Python", "Data Engineering", "Pandas", "DuckDB", "Polars", "Performance"]
license: all-rights-reserved
word_count: 268
reading_minutes: 1
citation: "Eddie's Blog. \"Pandas Should Go Extinct.\" 12 Sept 2026. https://eddie.codes/posts/pandas-should-go-extinct/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Pandas Should Go Extinct

> Eddie argues that Pandas forces data practitioners to adopt distributed compute infrastructure long before their data sizes justify it, and that DuckDB and Polars can fill the gap for workloads up to around 100 GB on a single machine. Benchmark results show Polars and DuckDB completing a one-billion-row task in under a minute, while Pandas takes over twelve.

> **Indexed summary.** This entry is an agent-written synopsis of an article first published at [eddie.codes](https://eddie.codes/posts/pandas-should-go-extinct/). Read the original for the full text.

The post opens with what the author calls the "Pandas cliff": users outgrow Pandas' in-memory model somewhere in the tens-of-gigabytes range and are typically pushed toward Spark, Databricks, or Snowflake — expensive, operationally complex systems designed for genuinely large-scale workloads most data teams will never reach.

The key evidence comes from Amazon's 2024 analysis of the Redshift fleet: roughly 95% of tables contain under 100 GB and about 87% of queries touch 80 GB or less. The author's conclusion is that most data work is a "Medium Data" problem solvable on a single machine with modern tooling.

## Key points

- Polars uses lazy evaluation, a query optimiser, and chunk-wise multi-threaded processing, achieving a 17× speed improvement over Pandas on the one-billion-row CSV benchmark.
- DuckDB achieves similar speed with just 546 MB of peak memory, compared to Pandas' 38 GB on the same task.
- Both tools support Apache Arrow as an interchange format, enabling incremental migration without rewriting existing Pandas code.
- On laptop hardware (16 GB RAM), Pandas forced 21 GB of swap; Polars and DuckDB completed the benchmark without touching swap.
- The author recommends DuckDB when SQL is natural or memory is the primary constraint, and Polars when a DataFrame API is preferred.

## Why it matters

The piece makes a concrete, data-backed case that the default upgrade path from Pandas — reaching for distributed systems — is frequently unnecessary and expensive. Pointing analysts toward single-machine alternatives has real cost and complexity implications for data teams.

---

*Source: [Pandas Should Go Extinct](https://eddie.codes/posts/pandas-should-go-extinct/)*
