TL;DR
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Polars 2.0 makes the streaming engine the default for LazyFrame collection and enables initial out-of-core spilling, aiming to reduce memory pressure on supported operations. The company also expands SQL support and reports strong results in its own TPC-H and TPC-DS tests; the benchmark claims have not been independently verified here.
Polars has released version 2.0, making its streaming engine the default when users call collect on a LazyFrame and enabling initial spill-to-disk support. The update could help data workloads that exceed available memory, but it also changes a key behavior: some operations no longer preserve row order by default.
Polars says the new default applies to LazyFrame queries. Its streaming engine processes data in batches rather than requiring the full query result to be held in memory, according to the project. In Polars 2.0, however, operations including joins, group-by, and unpivot do not guarantee observable row order unless users request it with maintain_order=True. Teams whose downstream work depends on a particular order may need to update queries or check their results after upgrading.
The release also turns on out-of-core processing, which lets certain operations spill data to disk as memory use grows. Polars says spilling begins at about 80% of RAM and has a default disk budget of 64 GB; those settings may need tuning. Current support includes sorts, window functions, and many expressions. Joins and group-by operations are not yet listed as supported for out-of-core execution, although the company says it plans to add them.
Other changes include expanded SQL support, optimizer and engine work, a new Map data type, and stricter handling of data types and explicitness. Polars identifies join reordering, common-subplan elimination, and dynamic predicates or bloom filters among the performance improvements. The announcement says the stricter rules are intended to provide faster feedback during development, including when iterating on AI-assisted code; it does not provide separate measurements for that claim.
Memory Limits and Row-Order Changes
The practical change for many existing users is not only the prospect of lower memory use, but the altered default execution behavior. Streaming can make larger queries more manageable, while disk spilling gives supported operations another way to complete when they cannot fit comfortably in RAM. That may reduce the need to restructure some workflows around available memory, though the release does not extend spill support to every operation.
The row-order caveat matters for code that assumes joins, group-bys, or unpivots will return rows in a predictable sequence. Users can request order preservation, but should test whether that option is needed and account for its cost in their workload. The SQL improvements may also make Polars more relevant to teams that want to use SQL rather than only its DataFrame APIs; the release announcement describes SQL as a first-class part of the project.
Polars also presents benchmark results as evidence that its engine can compete with established query engines. Those results are useful as a signal, but they come from the project’s own tests and depend on the hardware, data, queries, and scoring approach used. They should not be treated as a guarantee that Polars will be fastest on a reader’s production workload.
high-performance data processing laptop
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Release Tests Compare
For its SQL performance comparison, Polars says it tested queries generated from TPC-H and TPC-DS data against DuckDB 1.5.6, a DuckDB 2.0 alpha build, and DataFusion 54.0.0. Tests ran on two AWS machine types: a c7a.4xlarge with 16 vCPUs and 32 GB of memory, and a c7a.metal with 192 vCPUs and 384 GB. Each query ran five times in a hot setting; the report used the best of those runs and compared both total and geometric-mean query times.
According to Polars, its default configuration was fastest in all but one of the reported benchmarks. The project also says that restricting Polars to 32 cores was competitive or winning across the benchmarks. It reported that DataFusion timed out on TPC-DS query 72, timed out once on query 67, and ran out of memory on TPC-H query 18 on the smaller machine; those affected queries were excluded for all engines. Polars notes a scaling overhead on the 192-thread machine that hurt small queries and says it hopes to address the issue in a later release. It published a repository so others can replicate the tests.
Polars 2.0 also adds a native Map dtype corresponding to Arrow’s MapType. The project says that, before this release, Arrow maps were represented in Polars as lists of key-value structs. The new type provides map-oriented methods such as retrieving values by key and checking whether a key exists.
“Calling collect on a LazyFrame will now default to the streaming engine.”
— Polars, in its Polars 2.0 release announcement
As an affiliate, we earn on qualifying purchases.
Limits of Spill and Benchmark Claims
The release announcement does not specify a publication date, so the timing is described here as a release rather than assigned a calendar date. It also does not provide independent validation of the benchmark results. Polars shared its methodology and code repository, but results in other environments may differ; its tests used particular hardware, data formats, query sets, and a best-of-five scoring method.
The exact memory and performance effects of the new defaults will depend on query shape and data. Polars says spilling starts at roughly 80% of RAM and that the threshold may need tuning, but the announcement does not give a complete compatibility matrix or quantify gains across a broad set of real-world workloads. Support for out-of-core joins and group-bys remains a planned addition, not a feature confirmed as available in this release.
As an affiliate, we earn on qualifying purchases.
Testing the New Defaults
Users moving to Polars 2.0 should test representative lazy queries, with particular attention to whether they depend on row order after joins, group-bys, or unpivots. Where order is observable and required, the project points to maintain_order=True. Teams should also check disk-budget and memory settings against their environment, since spill behavior is enabled by default and may require tuning.
Polars says it plans to extend out-of-core support to joins and group-by operations and hopes to address the overhead it observed when using 192 threads in a future release. The project has made its benchmark repository available for replication. The source announcement does not give dates for either planned work, so their timing remains open.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main change in Polars 2.0?
LazyFrame collection now defaults to the streaming engine, and initial out-of-core spilling is enabled. Some operations do not preserve row order by default under the streaming engine.
Which operations can spill data to disk?
Polars lists sorts, window functions, and many expressions as currently supported. The project says out-of-core support for joins and group-by operations is planned, not yet available in this release.
How can users preserve row order?
For the operations covered by the release note, users can set maintain_order=True when they require observable row-order preservation. They should verify behavior in queries that depend on a particular sequence.
Did Polars independently prove it is faster than DuckDB and DataFusion?
No independent validation is described in the source. Polars ran and reported the benchmarks, using specified TPC-H and TPC-DS tests, hardware, and timing rules, and shared code intended to let others replicate the results.
What is the new Map dtype?
It is Polars’ direct support for Arrow MapType, representing key-value mappings. The release adds methods for operations such as looking up a key, checking for a key, and retrieving map keys or values.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
