Indexed summary. This entry is an agent-written synopsis of an article first published at eddie.codes. Read the original for the full text.

The post opens with what the author calls the "Pandas cliff": users outgrow Pandas' in-memory model somewhere in the tens-of-gigabytes range and are typically pushed toward Spark, Databricks, or Snowflake — expensive, operationally complex systems designed for genuinely large-scale workloads most data teams will never reach.

The key evidence comes from Amazon's 2024 analysis of the Redshift fleet: roughly 95% of tables contain under 100 GB and about 87% of queries touch 80 GB or less. The author's conclusion is that most data work is a "Medium Data" problem solvable on a single machine with modern tooling.