AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Polars says version 2.0 makes the streaming engine the default for LazyFrame collection and enables initial out-of-core processing that can spill supported operations to disk. The release also adds a Map dtype and expanded SQL support; the project’s benchmark results favor Polars in most tested cases, but reflect the developers’ own test setup.

Polars has shipped version 2.0, making its streaming engine the default when users collect a LazyFrame and enabling initial out-of-core processing that can spill supported query work to disk. The release also expands SQL support and adds a Map data type, while changing the default row-order guarantees for some operations, a compatibility issue users may need to address.

Under the new default, calling collect on a LazyFrame uses the streaming engine. Polars says this can bring memory and performance improvements for many queries. However, the engine does not preserve observable row order by default for certain operations, including joins, group-bys and unpivots. Users who need ordering for those operations can request it with maintain_order=True.

Version 2.0 also turns on an initial form of out-of-core processing, in which supported operations can spill data to disk as memory fills. The release post says spilling begins at about 80% of RAM, a threshold the developers say may need tuning, and the default disk budget is 64 GB. Sorts, window functions and many expressions are among the operations supported so far; joins and group-bys are not yet included.

Other changes include direct support for Arrow’s MapType as a Polars Map dtype, with dictionary-like operations such as retrieving values by key. Polars also says it has made SQL a first-class interface and improved its optimizer and engine, including join reordering, common-subplan elimination and dynamic predicates or bloom filters.

At a glance
announcementWhen: Announced in the Polars 2.0 release pos…
The developmentPolars has released version 2.0, shifting LazyFrame collection to streaming execution by default and enabling initial spill-to-disk support.

Memory Limits and Row-Order Changes

The new defaults affect both how workloads run and what users can rely on in their results. Streaming execution and disk spilling are intended to help queries handle larger workloads without requiring all intermediate data to remain in memory. The initial support is not universal, however: operations such as joins and group-bys cannot yet use the announced spill-to-disk path.

The row-order change is a practical migration concern. Users whose downstream code depends on the order of rows after joins, group-bys or unpivots may need to specify maintain_order=True and check their results. Polars’ release post presents the change as a trade-off associated with using streaming by default, rather than claiming that every query will become faster or use less memory.

Expanded SQL support may also make Polars more relevant to teams that want to run analytical queries through SQL as well as its existing interfaces. Polars reports favorable performance in its benchmarks, but the results should be read alongside the stated hardware, workload and test method rather than as a guarantee for every dataset or deployment.

Amazon

high performance data analysis laptop

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the SQL Benchmarks Were Run

Polars compared its SQL engine with DuckDB 1.5.6, DuckDB 2.0 alpha and DataFusion 54.0.0 on queries based on TPC-H and TPC-DS. Tests ran on two machines: one with 16 virtual CPUs and 32 GB of memory, and another with 192 virtual CPUs and 384 GB. The team ran each query five times in a hot setting and compared both total runtime and the geometric mean of query runtimes.

According to Polars, its default configuration was fastest on all but one of the reported benchmarks. The post says Polars and both DuckDB versions completed all queries, while DataFusion timed out on TPC-DS query 72, timed out once on query 67 and ran out of memory on TPC-H query 18 on the smaller machine. Those affected queries were excluded from the results for all engines. Polars also reports a fixed overhead when scaling to 192 threads that hurt smaller queries, and says a 32-core configuration was competitive or winning in all benchmarks.

The benchmark was conducted and published by the Polars team, which also provided a repository for replication. The release post describes the test data, query generation, storage and run procedure; these are project-reported results, not an independent evaluation. Polars says the 2.0 version bump had been explained in an earlier announcement, while this release post focuses on the features included in the shipped software.

“Calling collect on a LazyFrame will now default to the streaming engine.”

— Polars, in its Polars 2.0 release post

Amazon

large RAM external SSD for data processing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Initial Release

The release information does not establish how much faster or less memory-intensive Polars 2.0 will be for a particular user’s workload. The developers describe the gains as applying to many queries, but performance depends on factors such as query shape, data size and hardware. Their benchmark report also notes that scaling to 192 threads imposed overhead on smaller queries.

Out-of-core coverage remains partial: joins and group-bys are not supported by the initial spill-to-disk implementation, although Polars says it plans to add them. The release post says the spilling threshold is about 80% of RAM and may need tuning, but does not provide a date for broader support or a revised threshold. The source also does not specify a publication date for the release announcement.

Amazon

SQL database management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Broader Spill Support Ahead

Polars says it plans to extend out-of-core processing to joins and group-bys, which would broaden the range of queries able to spill to disk. No delivery date is given in the release post. The team also says it has diagnosed the large-machine scaling overhead and hopes to address it in the next release; this is a stated intention, not a confirmed schedule.

For users moving to version 2.0, the immediate next step is to check workloads that rely on row order and opt into order preservation where needed. Teams evaluating the SQL performance claims can consult the benchmark repository shared by Polars and run comparable tests against their own queries and infrastructure.

Amazon

out-of-core data processing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main change in Polars 2.0?

LazyFrame collection defaults to the streaming engine, and initial out-of-core processing is enabled so some operations can spill data to disk.

Does Polars 2.0 preserve row order by default?

Not for certain streaming operations, including joins, group-bys and unpivots. Users who need observable ordering can set maintain_order=True for supported operations.

Which operations can spill to disk?

The release post lists sorts, window functions and many expressions as supported. Joins and group-bys are not supported by the initial out-of-core implementation.

Did Polars prove it is faster than DuckDB and DataFusion?

No independent proof is provided. Polars reports that its default configuration was fastest on all but one of its TPC-H and TPC-DS benchmarks, but those are tests run and reported by Polars under specified hardware and query conditions.

What is the new Map dtype for?

It represents Arrow MapType directly in Polars and supports dictionary-like data, including key lookups and methods to check keys or retrieve keys and values.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microsoft Comic Chat Is Now Open Source

Microsoft has released Comic Chat as open source, allowing developers to access and modify the chat application code for the first time.

Construct a Lead Qualification System That Works While You’re Sleeping

Learn how to create an automated lead qualification system that filters high-quality prospects effortlessly, saving you time and boosting conversions.

Upcoming breaking changes for npm v12

npm v12 will introduce security-focused default changes, blocking scripts and dependencies unless explicitly allowed, with release expected in July 2026.

Keep Solo Performances On Track With Simple One-Page Run Sheets

A new tool automates show-day planning for solo performers using single-page run sheets, reducing errors and streamlining event prep.