{"article":{"slug":"release-of-polars-2-0","title":"Release of Polars 2.0","subtitle":null,"summary":"The Polars team ships Polars 2.0: initial out-of-core spill-to-disk support on by default, core engine and optimizer speedups, first-class SQL that leads DuckDB and DataFusion on TPC-H and TPC-DS style benchmarks, a new Map dtype, and stricter dtype handling, plus a detailed benchmark appendix.","content_type":"announcement","language":"en","canonical_url":"https://pola.rs/posts/release-polars-2/","author":{"name":null,"url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Polars","url":"https://pola.rs/","listing_slug":null,"listing":null},"topics":[{"name":"Data Engineering","slug":"data-engineering","url":"https://listedarticles.com/topics/data-engineering"},{"name":"Databases","slug":"databases","url":"https://listedarticles.com/topics/databases"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Announcements","slug":"announcements","url":"https://listedarticles.com/topics/announcements"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1287,"reading_minutes":6,"published_at":"2026-10-06T00:00:00.000Z","added_at":"2026-10-06T14:15:13.036Z","updated_at":"2026-10-06T14:15:13.036Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/release-of-polars-2-0","markdown_url":"https://listedarticles.com/articles/release-of-polars-2-0.md","example":false,"citation":"Polars. \"Release of Polars 2.0.\" 6 Oct 2026. https://pola.rs/posts/release-polars-2/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://pola.rs/posts/release-polars-2/"},"body_markdown":"Today we are shipping Polars 2.0. In the earlier announcement post we went through the rationale of the version bump. This post we will discuss what features 2.0 brings. Even though we didn’t intend to make it a big feature release, it still packs a lot to get enthousiastic about.\n\nLet’s go through the highlights of this release:\n\n- our initial version of out-of-core (spill-to-disk) support is enabled,\n- a lot of very core performance improvements,\n- first class SQL support, which together with the performance improvements has **Polars leading DataFusion and DuckDB in TPC-H and TPC-DS<sup>[1](https://pola.rs#user-content-fn-1)</sup> benchmarks** ,\n- a new `Map` dtype, and\n- stricter Polars on dtypes and explicitness, leading to faster feedback, and faster AI iteration.\n\n## Performance and SQL as a first class citizen\n\nPolars 2.0 will be the marking point where we will treat SQL as first class citizen. Polars SQL coverage has increased dramatically over last few months. We know we have been building a solid engine for the last couple of years. In Polars 2.0, we want to enable that to more workloads, including SQL. To make this performant, we shipped many improvements to our optimizer and engine. The highlights here join reordering, much better common-subplan-elimination and dynamic predicates/bloom filters.\n\nTo see how we perform on typical SQL benchmarks, we ran Polars SQL on data derived from TPC-H and TPC-DS<sup>[1](https://pola.rs#user-content-fn-1)</sup> and ran it against the latest DuckDB release (1.5.6), DuckDB 2.0 alpha (2.0.0.dev2610011535) and the latest DataFusion release (54.0.0) on a c7a.4xlarge (16 vCPUs, 32GB RAM) and a c7a.metal (192 vCPUs, 384GB RAM). Every query ran 5 times in a hot setting, with a separate process per query and a 60 second timeout. The file cache was cleared between each engine/benchmark (not between queries). For every query we take the best of the 5 runs, and we compare engines on both the sum and the geometric mean of those query times.\n\nThe data is generated with `tpcgen-cli parquet` compiled from source on commit `99bedae`. We looked at the default row-group sizes of `tpcgen-cli` and confirmed they are roughly similar to what Polars `scan_csv` piped through `sink_parquet` and Duckdb `COPY` produce. The SQL queries were generated with DuckDB 1.5.6’s `tpch_queries()` and `tpcds_queries()`. The data was stored on EBS.\n\nThe charts below show the runtime of each engine in seconds (lower is better), split by machine.\n\n**c7a.4xlarge (16 vCPUs, 32 GB)**\n\n**c7a.metal (192 vCPUs, 384 GB)**\n\nPolars and both DuckDB versions completed all queries. DataFusion timed out on TPC-DS q72 (and once on q67) and ran out of memory on TPC-H q18 on c7a.4xlarge; those queries are excluded from the results above for all engines.\n\nWe observe that default Polars is fastest on all but one benchmarks. Polars has a constant overhead when we scale to 192 threads, which hurts small data queries. In fact we see that Polars limited to 32 cores is competitive or winning in all benchmarks. We have diagnosed the cause on our end and will hopefully fix this problem in the next release. More information on the benchmarks can be found in the [appendix](https://pola.rs#benchmark-appendix). We encourage you to replicate our results and have shared a repository for this benchmark here: [https://github.com/pola-rs/polars-2.0-benchmark](https://github.com/pola-rs/polars-2.0-benchmark).\n\n## Streaming engine and OOC as default\n\nThis is the one of the biggest impact changes of 2.0. Calling `collect` on a `LazyFrame` will now default to the streaming engine, leading to massive memory and performance improvements on most queries. The reason this required a major version bump is that the streaming engine doesn’t guarantee row-order by default for certain operations (`join`, `group_by`, `unpivot`, etc.). If you require observable row-order in those operations, you can opt in to that by setting `maintain_order=True`.\n\nOut-of-core (spill to disk) is now enabled by default. It starts spilling at ~80% of RAM (this may need tuning). Operations that support out-of-core at this moment (sort, window functions, many expressions) can now start spilling to disk to finish a query. The default disk budget is 64GB. In the coming time we will enable out-of-core for joins and group-by’s as well.\n\nThese two changes will make Polars much more resillient in high-memory workloads for casual data practicioners. And with out-of-core `join` and `group-by` on our roadmap, this resilience will improve even more.\n\n## New Map datatype\n\nPolars now supports the Arrow `MapType` directly as a Polars `Map` dtype. You can think of a `Map` as a Python dictionary, mapping keys to values. Before 2.0 the Arrow `MapType` was read in Polars as `List(Struct({\"key\": ..., \"value\": ...})).`\n\n```\ndf = pl.DataFrame(\n    {\n        \"user\": [\"alice\", \"bob\", \"carol\"],\n        \"scores\": pl.Series(\n            [{\"math\": 90, \"art\": 75}, {\"math\": 60}, {}],\n            dtype=pl.Map(pl.String, pl.Int64),\n        ),\n        \"subject\": [\"art\", \"art\", \"math\"],\n    }\n)\n```\n```\nshape: (3, 3)\n┌───────┬─────────────────────────┬─────────┐\n│ user  ┆ scores                  ┆ subject │\n│ ---   ┆ ---                     ┆ ---     │\n│ str   ┆ map[str, i64]           ┆ str     │\n╞═══════╪═════════════════════════╪═════════╡\n│ alice ┆ {\"math\": 90, \"art\": 75} ┆ art     │\n│ bob   ┆ {\"math\": 60}            ┆ art     │\n│ carol ┆ {}                      ┆ math    │\n└───────┴─────────────────────────┴─────────┘\n```\n```\n# Key lookups and dictionary-like methods:\ndf.select(\n    \"user\",\n    pl.col(\"scores\").map.get(\"math\").alias(\"math\"),                # fixed key\n    pl.col(\"scores\").map.get(pl.col(\"subject\")).alias(\"by_subject\"),  # key from another column\n    pl.col(\"scores\").map.contains_key(\"art\").alias(\"has_art\"),\n    pl.col(\"scores\").map.len().alias(\"n\"),\n    pl.col(\"scores\").map.keys().alias(\"keys\"),\n    pl.col(\"scores\").map.values().alias(\"values\"),\n)\n```\n```\n┌───────┬──────┬────────────┬─────────┬─────┬─────────────────┬───────────┐\n│ user  ┆ math ┆ by_subject ┆ has_art ┆ n   ┆ keys            ┆ values    │\n│ ---   ┆ ---  ┆ ---        ┆ ---     ┆ --- ┆ ---             ┆ ---       │\n│ str   ┆ i64  ┆ i64        ┆ bool    ┆ u32 ┆ list[str]       ┆ list[i64] │\n╞═══════╪══════╪════════════╪═════════╪═════╪═════════════════╪═══════════╡\n│ alice ┆ 90   ┆ 75         ┆ true    ┆ 2   ┆ [\"math\", \"art\"] ┆ [90, 75]  │\n│ bob   ┆ 60   ┆ null       ┆ false   ┆ 1   ┆ [\"math\"]        ┆ [60]      │\n│ carol ┆ null ┆ null       ┆ false   ┆ 0   ┆ []              ┆ []        │\n└───────┴──────┴────────────┴─────────┴─────┴─────────────────┴───────────┘\n```\nAs a supported dtype, the map type will now have dedicated expressions, like key lookups, iteration over values and other dictionary like methods.\n\n## Stricter Polars\n\nPolars aims to be strict and fail fast. Errors should ideally raise up-front, not 20 minutes into a pipeline. Implicit behavior on data-mismatches should be opt-in, not a default, since those mismatches can hide bugs. This strictness has become even more valuable with the rise of AI-driven development. Agents can validate a query’s structure early by calling `collect_schema()`, which resolves types and catches schema-level mismatches without materializing any data. This ensures fast feedback, meaning agents and humans can iterate faster. Not all errors can be caught during compilation of the query plan, some depend on data. In these cases Polars defaults to stricter behavior to ensure inconsistencies are caught instead of silently producing different results. See [previous posts](https://pola.rs/posts/announcing-polars-2/)) with some examples in where Polars has gotten more strict.\n\n## Last words\n\nWe are very excited that Polars 2.0 is out. Coming months we’ll improve on the road were in. Better out-of-core, better scaling at large CPU-counts and on Polars Cloud we aim to be the fastest distributed engine available. We also started working on GeoPolars and hope to deliver more news on this soon.\nIf you find any problem with our new release, please open an issue: [https://github.com/pola-rs/polars/issues](https://github.com/pola-rs/polars/issues). And finally, to help you with upgrading to 2.0, we have posted [migration guide](https://docs.pola.rs/releases/upgrade/2/).\n\n## Benchmark Appendix\n\nThe absolute numbers (in seconds) are below. The bold number is the fastest engine in each row; the darker the shading, the slower an engine is compared to the fastest. Polars with 32 threads was only run on c7a.metal.\n\n**Sum of query times**\n\n**Geometric mean of query times**\n\nPolars also scales well with more cores on larger data. Moving from 16 to 192 vCPUs at SF100 makes Polars 3.8x faster on TPC-H and 2.2x faster on TPC-DS (by sum), compared to 3.2x and 1.9x for DuckDB 1.5.6, 2.2x and 1.5x for the DuckDB 2.0 alpha, and 1.7x and 1.0x for DataFusion. At SF10 the extra cores don’t help Polars with its default settings: it is equally fast on TPC-H and 1.8x slower on TPC-DS, while DuckDB 1.5.6 still gets 1.8x and 1.3x faster. Polars with 32 threads only ran on the c7a.metal, so it is not part of this comparison.\n","body_html":"<p>Today we are shipping Polars 2.0. In the earlier announcement post we went through the rationale of the version bump. This post we will discuss what features 2.0 brings. Even though we didn’t intend to make it a big feature release, it still packs a lot to get enthousiastic about.</p>\n<p>Let’s go through the highlights of this release:</p>\n<ul><li>our initial version of out-of-core (spill-to-disk) support is enabled,</li><li>a lot of very core performance improvements,</li><li>first class SQL support, which together with the performance improvements has <strong>Polars leading DataFusion and DuckDB in TPC-H and TPC-DS&lt;sup&gt;<a href=\"https://pola.rs#user-content-fn-1\" rel=\"nofollow ugc noopener\">1</a>&lt;/sup&gt; benchmarks</strong> ,</li><li>a new <code>Map</code> dtype, and</li><li>stricter Polars on dtypes and explicitness, leading to faster feedback, and faster AI iteration.</li></ul>\n<h2 id=\"performance-and-sql-as-a-first-class-citizen\">Performance and SQL as a first class citizen</h2>\n<p>Polars 2.0 will be the marking point where we will treat SQL as first class citizen. Polars SQL coverage has increased dramatically over last few months. We know we have been building a solid engine for the last couple of years. In Polars 2.0, we want to enable that to more workloads, including SQL. To make this performant, we shipped many improvements to our optimizer and engine. The highlights here join reordering, much better common-subplan-elimination and dynamic predicates/bloom filters.</p>\n<p>To see how we perform on typical SQL benchmarks, we ran Polars SQL on data derived from TPC-H and TPC-DS&lt;sup&gt;<a href=\"https://pola.rs#user-content-fn-1\" rel=\"nofollow ugc noopener\">1</a>&lt;/sup&gt; and ran it against the latest DuckDB release (1.5.6), DuckDB 2.0 alpha (2.0.0.dev2610011535) and the latest DataFusion release (54.0.0) on a c7a.4xlarge (16 vCPUs, 32GB RAM) and a c7a.metal (192 vCPUs, 384GB RAM). Every query ran 5 times in a hot setting, with a separate process per query and a 60 second timeout. The file cache was cleared between each engine/benchmark (not between queries). For every query we take the best of the 5 runs, and we compare engines on both the sum and the geometric mean of those query times.</p>\n<p>The data is generated with <code>tpcgen-cli parquet</code> compiled from source on commit <code>99bedae</code>. We looked at the default row-group sizes of <code>tpcgen-cli</code> and confirmed they are roughly similar to what Polars <code>scan_csv</code> piped through <code>sink_parquet</code> and Duckdb <code>COPY</code> produce. The SQL queries were generated with DuckDB 1.5.6’s <code>tpch_queries()</code> and <code>tpcds_queries()</code>. The data was stored on EBS.</p>\n<p>The charts below show the runtime of each engine in seconds (lower is better), split by machine.</p>\n<p><strong>c7a.4xlarge (16 vCPUs, 32 GB)</strong></p>\n<p><strong>c7a.metal (192 vCPUs, 384 GB)</strong></p>\n<p>Polars and both DuckDB versions completed all queries. DataFusion timed out on TPC-DS q72 (and once on q67) and ran out of memory on TPC-H q18 on c7a.4xlarge; those queries are excluded from the results above for all engines.</p>\n<p>We observe that default Polars is fastest on all but one benchmarks. Polars has a constant overhead when we scale to 192 threads, which hurts small data queries. In fact we see that Polars limited to 32 cores is competitive or winning in all benchmarks. We have diagnosed the cause on our end and will hopefully fix this problem in the next release. More information on the benchmarks can be found in the <a href=\"https://pola.rs#benchmark-appendix\" rel=\"nofollow ugc noopener\">appendix</a>. We encourage you to replicate our results and have shared a repository for this benchmark here: <a href=\"https://github.com/pola-rs/polars-2.0-benchmark\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/pola-rs/polars-2.0-benchmark\" rel=\"nofollow ugc noopener\">https://github.com/pola-rs/polars-2.0-benchmark</a></a>.</p>\n<h2 id=\"streaming-engine-and-ooc-as-default\">Streaming engine and OOC as default</h2>\n<p>This is the one of the biggest impact changes of 2.0. Calling <code>collect</code> on a <code>LazyFrame</code> will now default to the streaming engine, leading to massive memory and performance improvements on most queries. The reason this required a major version bump is that the streaming engine doesn’t guarantee row-order by default for certain operations (<code>join</code>, <code>group_by</code>, <code>unpivot</code>, etc.). If you require observable row-order in those operations, you can opt in to that by setting <code>maintain_order=True</code>.</p>\n<p>Out-of-core (spill to disk) is now enabled by default. It starts spilling at ~80% of RAM (this may need tuning). Operations that support out-of-core at this moment (sort, window functions, many expressions) can now start spilling to disk to finish a query. The default disk budget is 64GB. In the coming time we will enable out-of-core for joins and group-by’s as well.</p>\n<p>These two changes will make Polars much more resillient in high-memory workloads for casual data practicioners. And with out-of-core <code>join</code> and <code>group-by</code> on our roadmap, this resilience will improve even more.</p>\n<h2 id=\"new-map-datatype\">New Map datatype</h2>\n<p>Polars now supports the Arrow <code>MapType</code> directly as a Polars <code>Map</code> dtype. You can think of a <code>Map</code> as a Python dictionary, mapping keys to values. Before 2.0 the Arrow <code>MapType</code> was read in Polars as <code>List(Struct({&quot;key&quot;: ..., &quot;value&quot;: ...})).</code></p>\n<pre><code>df = pl.DataFrame(\n    {\n        &quot;user&quot;: [&quot;alice&quot;, &quot;bob&quot;, &quot;carol&quot;],\n        &quot;scores&quot;: pl.Series(\n            [{&quot;math&quot;: 90, &quot;art&quot;: 75}, {&quot;math&quot;: 60}, {}],\n            dtype=pl.Map(pl.String, pl.Int64),\n        ),\n        &quot;subject&quot;: [&quot;art&quot;, &quot;art&quot;, &quot;math&quot;],\n    }\n)</code></pre>\n<pre><code>shape: (3, 3)\n┌───────┬─────────────────────────┬─────────┐\n│ user  ┆ scores                  ┆ subject │\n│ ---   ┆ ---                     ┆ ---     │\n│ str   ┆ map[str, i64]           ┆ str     │\n╞═══════╪═════════════════════════╪═════════╡\n│ alice ┆ {&quot;math&quot;: 90, &quot;art&quot;: 75} ┆ art     │\n│ bob   ┆ {&quot;math&quot;: 60}            ┆ art     │\n│ carol ┆ {}                      ┆ math    │\n└───────┴─────────────────────────┴─────────┘</code></pre>\n<pre><code># Key lookups and dictionary-like methods:\ndf.select(\n    &quot;user&quot;,\n    pl.col(&quot;scores&quot;).map.get(&quot;math&quot;).alias(&quot;math&quot;),                # fixed key\n    pl.col(&quot;scores&quot;).map.get(pl.col(&quot;subject&quot;)).alias(&quot;by_subject&quot;),  # key from another column\n    pl.col(&quot;scores&quot;).map.contains_key(&quot;art&quot;).alias(&quot;has_art&quot;),\n    pl.col(&quot;scores&quot;).map.len().alias(&quot;n&quot;),\n    pl.col(&quot;scores&quot;).map.keys().alias(&quot;keys&quot;),\n    pl.col(&quot;scores&quot;).map.values().alias(&quot;values&quot;),\n)</code></pre>\n<pre><code>┌───────┬──────┬────────────┬─────────┬─────┬─────────────────┬───────────┐\n│ user  ┆ math ┆ by_subject ┆ has_art ┆ n   ┆ keys            ┆ values    │\n│ ---   ┆ ---  ┆ ---        ┆ ---     ┆ --- ┆ ---             ┆ ---       │\n│ str   ┆ i64  ┆ i64        ┆ bool    ┆ u32 ┆ list[str]       ┆ list[i64] │\n╞═══════╪══════╪════════════╪═════════╪═════╪═════════════════╪═══════════╡\n│ alice ┆ 90   ┆ 75         ┆ true    ┆ 2   ┆ [&quot;math&quot;, &quot;art&quot;] ┆ [90, 75]  │\n│ bob   ┆ 60   ┆ null       ┆ false   ┆ 1   ┆ [&quot;math&quot;]        ┆ [60]      │\n│ carol ┆ null ┆ null       ┆ false   ┆ 0   ┆ []              ┆ []        │\n└───────┴──────┴────────────┴─────────┴─────┴─────────────────┴───────────┘</code></pre>\n<p>As a supported dtype, the map type will now have dedicated expressions, like key lookups, iteration over values and other dictionary like methods.</p>\n<h2 id=\"stricter-polars\">Stricter Polars</h2>\n<p>Polars aims to be strict and fail fast. Errors should ideally raise up-front, not 20 minutes into a pipeline. Implicit behavior on data-mismatches should be opt-in, not a default, since those mismatches can hide bugs. This strictness has become even more valuable with the rise of AI-driven development. Agents can validate a query’s structure early by calling <code>collect_schema()</code>, which resolves types and catches schema-level mismatches without materializing any data. This ensures fast feedback, meaning agents and humans can iterate faster. Not all errors can be caught during compilation of the query plan, some depend on data. In these cases Polars defaults to stricter behavior to ensure inconsistencies are caught instead of silently producing different results. See <a href=\"https://pola.rs/posts/announcing-polars-2/\" rel=\"nofollow ugc noopener\">previous posts</a>) with some examples in where Polars has gotten more strict.</p>\n<h2 id=\"last-words\">Last words</h2>\n<p>We are very excited that Polars 2.0 is out. Coming months we’ll improve on the road were in. Better out-of-core, better scaling at large CPU-counts and on Polars Cloud we aim to be the fastest distributed engine available. We also started working on GeoPolars and hope to deliver more news on this soon.\nIf you find any problem with our new release, please open an issue: <a href=\"https://github.com/pola-rs/polars/issues\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/pola-rs/polars/issues\" rel=\"nofollow ugc noopener\">https://github.com/pola-rs/polars/issues</a></a>. And finally, to help you with upgrading to 2.0, we have posted <a href=\"https://docs.pola.rs/releases/upgrade/2/\" rel=\"nofollow ugc noopener\">migration guide</a>.</p>\n<h2 id=\"benchmark-appendix\">Benchmark Appendix</h2>\n<p>The absolute numbers (in seconds) are below. The bold number is the fastest engine in each row; the darker the shading, the slower an engine is compared to the fastest. Polars with 32 threads was only run on c7a.metal.</p>\n<p><strong>Sum of query times</strong></p>\n<p><strong>Geometric mean of query times</strong></p>\n<p>Polars also scales well with more cores on larger data. Moving from 16 to 192 vCPUs at SF100 makes Polars 3.8x faster on TPC-H and 2.2x faster on TPC-DS (by sum), compared to 3.2x and 1.9x for DuckDB 1.5.6, 2.2x and 1.5x for the DuckDB 2.0 alpha, and 1.7x and 1.0x for DataFusion. At SF10 the extra cores don’t help Polars with its default settings: it is equally fast on TPC-H and 1.8x slower on TPC-DS, while DuckDB 1.5.6 still gets 1.8x and 1.3x faster. Polars with 32 threads only ran on the c7a.metal, so it is not part of this comparison.</p>","headings":[{"level":2,"text":"Performance and SQL as a first class citizen","id":"performance-and-sql-as-a-first-class-citizen"},{"level":2,"text":"Streaming engine and OOC as default","id":"streaming-engine-and-ooc-as-default"},{"level":2,"text":"New Map datatype","id":"new-map-datatype"},{"level":2,"text":"Stricter Polars","id":"stricter-polars"},{"level":2,"text":"Last words","id":"last-words"},{"level":2,"text":"Benchmark Appendix","id":"benchmark-appendix"}]}}