If your DuckDB pipelines read large Parquet or CSV files straight out of S3, the 2.0 alpha is worth piloting this week on non-critical work; if your lake is thousands of tiny files, the measured gain shrinks to almost nothing, and if your reads never leave local disk, these numbers say nothing about your runtimes. That is the honest summary of every performance number in this article: a set of author-reported benchmarks from MotherDuck’s “Why DuckDB 2.0 is faster” post, measured on a single M5 laptop over home internet to us-east-1, with enable_external_file_cache set to false so every run actually hit S3, both versions on the same machine, and the headline query reading one column of four. The author says so plainly: “Every number below is from one machine (an M5 laptop) and my home internet, which slows both versions about equally. Run your own before quoting them ;)”. DuckDB 2.0 itself is “coming this fall” per the same vendor post, which is a vendor-reported intent rather than a dated GA schedule. Within those limits, the numbers describe a clear workload shape where the upgrade pays, and an equally clear shape where it does not.
What async I/O actually changes
The mechanism is a separation of concerns that did not exist in 1.5. In DuckDB 1.5.5, per the post, “each of the 18 workers does both jobs in turn: download, wait, decode, download, wait. While a worker waits its CPU is idle, while it decodes it has no download in flight, and you never get more than 18 downloads going.” Network latency and CPU decode are serialized on the same threads, so a scan of remote Parquet is bounded by round trips even when the CPU has idle capacity.
In 2.0, a separate download-only thread pool fetches row groups ahead of the workers, so the network can stay saturated while decode keeps chewing. One setting governs it: “One setting drives this: read_ahead_depth, how many row groups the download pool may fetch ahead of the workers. It defaults to -1 (automatic, sized from your thread count), so async I/O is on out of the box. Set it to 0 and you get the 1.5 behaviour back.” Two practical consequences follow. You do not need to change a single query to get the speedup. And you have a clean rollback switch: if the alpha misbehaves, read_ahead_depth = 0 restores the old I/O pattern without a downgrade.
This design also explains why the gains should be largest where network waits dominate. An independent study of DuckDB running on xNVMe makes the same point from the storage side: async I/O pays only when concurrency design allows many reads in flight, and its authors characterize I/O-interface-level improvements for local storage as not significantly high. Treat the S3 numbers as evidence about remote reads, not as a promise about your NVMe.
The numbers, measured exactly as described
All four scenarios below are author-reported from that one laptop, external file cache disabled, home internet to us-east-1.
| S3 read (author-reported, same laptop) | DuckDB 1.5.5 | DuckDB 2.0 alpha | Speedup |
|---|---|---|---|
| One 2.2 GB Parquet file, one column of four | 18.8 s | 7.7 s | ~2.4× |
| 23 large Parquet files, 13.6 GB, one column | 11.8 s | 3.9 s | ~3.0× |
| One 1.7 GB plain CSV | 116 s | 55 s | ~2.1× |
| 30 tiny Parquet files, about 1 MB each | 3.7 s | 3.3 s | ~1.1× |
The headline query counted votes per type in a 2.2 GB Stack Overflow votes file (228 million rows, 2,268 row groups, per the vendor’s post) while reading only one of four columns, roughly 230 MB of the file. Three things stand out.
First, the gain is not Parquet-specific. The plain CSV went from 116 s to 55 s, about 2.1×, per the same vendor table, because the bottleneck being removed is the network wait, not the format decode. If your lake still has large CSV exports sitting in a bucket, this release treats them about as kindly as Parquet for remote reads.
Second, many large files beat one large file in relative terms: 11.8 s down to 3.9 s across 23 files totaling 13.6 GB. With more row groups to prefetch, the download pool has more useful work to interleave.
Third, the absolute numbers will look different from cloud compute. The author adds: “Quick caveat: as I said in the introduction, this is my home internet to us-east-1, so both numbers are slow. Expect things to go faster if you run this from cloud compute.” A faster link shrinks the network wait, which is precisely the cost async I/O exists to hide, so the 2–3× factor could compress on a well-connected EC2 instance in the same region as the bucket. The direction of the result is credible; the multiplier is not portable. That is an inference from the mechanism, not a measured finding.
Where the speedup runs out
The most useful result in the post is the one the vendor headline will not lead with. On 30 Parquet files of about 1 MB each, 2.0 alpha moved 3.7 s to 3.3 s. The author’s explanation: with tiny files the cost is “per-file round trips (footer, then data) that reading ahead cannot remove.” Every Parquet file costs a footer read before any data read; with tiny files the workload is dominated by those fixed per-file round trips, and prefetching row groups does not shorten them.
This matters because “my data is on S3” is not a workload shape. A compacted lake of multi-hundred-megabyte files and a streaming ingest that drops one file per minute are both “on S3,” and this release helps only one of them. If your file inventory skews toward the megabyte range, the fix is compaction, not a version upgrade.
The second boundary is local disk. The post’s own summary is that “reading data over S3 is 2x to 3x faster in 2.0 with zero query changes,” and its async I/O benchmarks target S3; the xNVMe study above suggests local I/O gains from interface-level async work are modest. If your pipelines read from attached NVMe, the post’s S3 numbers say nothing about your runtimes, and I would not assume the async I/O change moves them.
Row-group shape still rules, whatever the version
The upgrade cannot rescue a badly shaped file. DuckDB’s own writer defaults to 122,880 rows per row group, per a SmartNIC Parquet-decoding study from ETH Zürich and Apple, while TPC-H files written with Arrow’s Parquet writer use groups of 2^20 rows, roughly 8.5× larger, and those files “achieve worse overall query throughput in DuckDB” because larger groups leave far less parallelism for decoding and execution. Decode parallelism is the very thing async I/O is designed to feed; if your files arrive from an Arrow-based writer with giant row groups, 2.0’s download pool has less to overlap, and no engine release fixes that. Rewriting the files does.
The same paper documents that DuckDB’s virtual file systems coalesce column-chunk reads based on a tunable maximum byte gap between reads, which is the mechanism relevant to wide scans. The headline benchmark read one narrow column of four, and the post’s other large-file Parquet row reads one column too. Wide scans, where coalescing decides how the I/O is issued, sit outside that evidence. Analysts running SELECT *-shaped scans should treat their case as unmeasured and test it themselves.
The cache the benchmark switched off
Every number above was run with enable_external_file_cache = false (the post sets it with the comment “so every run really hits S3”), so each query paid full S3 retrieval cost. Leave the cache on, and the economics of repeat queries change: the second run against a warm cache is a local read, where the async I/O gains described here do not apply. Which files deserve cache residency is a genuinely hard problem; the classic on-line file caching analysis formalizes it as deciding “which files to keep in a fixed-size cache so as to minimize the sum of the retrieval costs for files that are not in the cache when requested,” and its Landlord algorithm, a generalization of LRU, carries an optimal guarantee among deterministic on-line strategies. The practical point is simpler than the theory: if your workload re-queries the same remote files daily, caching may already be absorbing much of the cost that async I/O targets, and your 2.0 delta will be smaller than the table suggests. If every query touches fresh data, the benchmark’s cold-cache setup matches yours.
Recursive CTEs: the other measured win
The recursive-CTE rewrite is described mechanistically, with the table read once and a parent-column lookup built once so cost scales with rows actually touched rather than rounds times table size, and this one comes with timings: on an ancestry walk over 20,000 commits, the author reports 1.5.5 at “1.8 s to 16 s across runs” and the 2.0 alpha at “0.10 s, every run.” The “40x on graph reachability” figure is the DuckDB team’s claim, relayed by the post rather than measured by its author. If recursive CTEs are your pain point, there is a number to plan around.
Alpha risk: what pins you to 1.5.x
An alpha earns its name. The post ships performance numbers, not stability data, community soak time, or an extension-compatibility matrix for 2.0. Extension coupling is a real failure mode in the DuckDB ecosystem: MobilityDuck, a third-party mobility-data extension, currently ships without Windows support “due to compatibility limitations in the underlying MEOS library,” a 1.x-era example of how an extension’s dependencies can strand a platform. That example illustrates the coupling risk, not a known 2.0 regression, but it is the shape of the thing that breaks when you upgrade a production stack before its extensions have caught up.
So who upgrades when
Here is how I would map the evidence to a decision:
- Pilot the alpha now if your slow queries are large-file scans against object storage, the work can tolerate a restart, and you can afford to re-run your own benchmarks. Async I/O is default-on, zero query changes are required, and
read_ahead_depth = 0is an instant rollback of the new I/O path. - Wait for GA if this is production, if you depend on third-party extensions, or if a failed pipeline rebuilds a customer’s work. The vendor-reported “coming this fall” window makes the wait short, and the alpha’s 2–3× author-reported gains are not worth an unquantified stability risk on workloads that have to run every night.
- Skip the upgrade as a performance fix if your lake is mostly ~1 MB files or your reads are local. Compaction and file layout will return more than any release notes will.
- Regardless of version, check who writes your Parquet. If an Arrow-based writer is producing ~2^20-row groups, reshaping those files is worth more than the engine upgrade.
Whichever branch you take, run your own scan on your own bucket before quoting a speedup to anyone. Every number in this article traces to one laptop on one home connection, and the author told you to.
Frequently Asked Questions
Does the async I/O speedup apply to local disk reads?
The post’s own summary is that “reading data over S3 is 2x to 3x faster in 2.0 with zero query changes,” and its async I/O benchmarks target S3; the xNVMe study above suggests local I/O gains from interface-level async work are modest. If your pipelines read from attached NVMe, the post’s S3 numbers say nothing about your runtimes, and I would not assume the async I/O change moves them.
Why does the speedup disappear for small Parquet files?
On 30 Parquet files of about 1 MB each, 2.0 alpha moved 3.7 s to 3.3 s. The author’s explanation: with tiny files the cost is “per-file round trips (footer, then data) that reading ahead cannot remove.” Every Parquet file costs a footer read before any data read; with tiny files the workload is dominated by those fixed per-file round trips, and prefetching row groups does not shorten them.

Join the discussion
Share a useful perspective or ask a question about this article.