DuckDB 2.0 alpha cut one measured S3 Parquet scan from 18.8 seconds to 7.7 seconds, according to MotherDuck’s benchmark post, and that number is real but far narrower than it looks. The measurement is one query, one file, one author’s laptop: a 2.2 GB Parquet file of Stack Overflow votes on S3, 228 million rows across 2,268 row groups, one column of four, about 230 MB (same post). It is author-reported, unreplicated, run against an alpha build rather than the current stable, and the author warns: “Every number below is from one machine (an M5 laptop) and my home internet, which slows both versions about equally. Run your own before quoting them ;)”
So the practical answer to “should we move S3 Parquet pipelines to DuckDB 2.0 now?” is: not into production, and not on the strength of that 2.4x figure. Keep production on stable 1.5.6 (or 1.4.5 LTS), and spend the few weeks before the fall stable release re-benchmarking your own queries against the 2.0.0-dev preview with the file cache disabled. Whether your workload captures the async I/O win depends on its shape, and this article lays out which shapes plausibly benefit, which probably don’t, and how to check against your own queries.
What DuckDB 2.0 actually ships, and what is still a promise
DuckDB’s official v2.0 preview post, published 2026-08-17, is direct about scope: “TL;DR: DuckDB v2.0 is coming this fall. In this post, we preview its headline features: DuckDB as a server, triggers, the VARIANT type, asynchronous I/O, a new SQL parser, a new storage format, and much more.” The release, codenamed Cyanoptera, is built from over 10,000 commits since v1.5 shipped in March (same post), and the same post warns that “v2.0 ships a new SQL parser, a new default storage format, a reworked C API, and a small number of carefully chosen breaking changes.”
That inventory matters as much as the speedup. A major version with a new default storage format and breaking changes is a migration project, not a version bump, and it prices the cost of moving early.
MotherDuck’s post is explicit about its own scope too: it covers the three features the author considers most important, with the author’s own numbers, plus a few hidden gems from the commit logs. Async I/O is the first; the second is the rewritten recursive CTE engine, where the author notes that “The DuckDB team rewrote the recursive CTE engine and claims 40x on graph reachability” (MotherDuck’s post). And as of 2026-10-11 the DuckDB install page lists 2.0.0-dev as a preview alongside 1.5.6 (current) and 1.4.5 (LTS), with the explicit warning: “The preview version is a development build of the next DuckDB release. It is constantly in flux and less suitable for production use than the current and LTS versions.”
The S3 measurements we have, and their exact shapes
The headline query, in the author’s words, “reads a 2.2 GB Parquet file on S3 (Stack Overflow votes, 228 million rows, 2268 row groups) and counts votes per type” (same post). It “reads one column out of four, about 230 MB” (same post).
That is the query behind the headline 18.8 s to 7.7 s result, but it is not the only measurement: the same post (MotherDuck) runs four S3 reads:
| S3 read | 1.5.5 | 2.0 alpha |
|---|---|---|
| One 2.2 GB Parquet file, one column | 18.8 s | 7.7 s |
| 23 large Parquet files, 13.6 GB, one column | 11.8 s | 3.9 s |
| One 1.7 GB plain CSV | 116 s | 55 s |
| 30 tiny Parquet files, about 1 MB each | 3.7 s | 3.3 s |
All four runs came from the same M5 laptop over the author’s home connection, with the external file cache disabled so every run genuinely hit S3. The second row shows the same narrow projection scaled across 23 large files, 13.6 GB total, and it also wins big: 11.8 s to 3.9 s (same table). The fourth row (3.7 s to 3.3 s) is the measured boundary of the win, and the author’s comment on it is mechanical: the time there is “per-file round trips (footer, then data) that reading ahead cannot remove.” In the author’s words, “Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.”
Three properties of the headline workload are doing the work, and each one points at a different reader:
- Narrow projection. One column of four means roughly 10% of the file’s bytes are needed (230 MB of 2.2 GB, per MotherDuck). The query is dominated by fetching many small, scattered byte ranges.
- Many row groups. 2,268 row groups in a 2.2 GB file averages under 1 MB per group, per the same post, so the useful bytes are spread across thousands of separate reads rather than a few large sequential ones.
- Remote storage. Each read pays S3 round-trip latency. On local NVMe the same access pattern would cost far less.
If your S3 queries look like the first two rows, the alpha’s results are at least directionally relevant to you. If they look like the fourth row, the post has already measured your answer: almost nothing changes.
Why async I/O rewards that shape
The mechanism is straightforward. A synchronous reader issues a request, waits for the response, then issues the next one. When a scan needs thousands of small byte ranges from a high-latency store, the engine spends most of the wall clock waiting on the network, not computing. Asynchronous I/O lets DuckDB keep many requests in flight at once, so latency overlaps with itself. The more, smaller, and more remote the reads, the more there is to overlap. In the alpha this is on by default: read_ahead_depth defaults to -1 (automatic), and setting it to 0 restores the 1.5 behavior (MotherDuck’s post).
This connects to how DuckDB already reads Parquet. Its table functions support projection pushdown and filter pushdown, and the Oasis paper on Parquet access costs documents the layout the DuckDB writer produces: it “emits row groups of 122,880 rows, splitting purely by row count with no byte-based limit” (same paper), and because “data pages are only split once they reach 100 MiB of uncompressed data” (same paper), at that row group size “each column chunk in practice contains a single data page.” The fragmentation async I/O can overlap therefore lives between row groups, not between pages inside one. The benchmark file is “cut into 2268 row groups of about 122,000 rows” (MotherDuck’s post), matching the writer’s 122,880-row default, and that means thousands of separate column-chunk fetches from S3. Row-group sizing, which your Parquet writer controls, directly changes how many separate reads a remote scan needs and how much a 2.0-style reader can recover.
There is also a reason to expect uneven results rather than a universal win. The DuckDB-on-xNVMe paper notes that DuckDB’s storage layer historically used the synchronous POSIX I/O interface and that “the design of DuckDB leads to data dependencies that require careful handling when introducing asynchronous I/Os.” Async I/O in this engine is an engineering-hard change with ordering constraints, not a free throughput multiplier applied everywhere. That is consistent with what MotherDuck’s table shows: large gains on big remote reads, and nothing to recover on per-file round trips.
One measured near-null case, and inference for the rest
The tiny-files row above settles one half of the “when it doesn’t” question with a measurement rather than an argument: async I/O is not a universal win, and per-file round trips (“footer, then data”) serialize no matter how aggressively the download pool reads ahead (MotherDuck’s post). The remaining shapes below are routing inference from mechanism, not measured counter-cases, and I am labeling them that way deliberately.
With that said, the mechanism implies three shapes where the async win should shrink or vanish:
- CPU-bound queries. If your query time is spent in expression evaluation, joins, or aggregation rather than waiting on S3, overlapping network requests buys nothing. DuckDB’s own preview post shows what a CPU-side win looks like instead: a recursive-CTE single-source reachability query over a 1M-edge graph runs in 4.90 s on v1.5.4 versus 0.12 s on the v2.0 preview (DuckDB’s preview post). That is a different mechanism (CTE execution improvements) for a different workload, and conflating its ~40x with the S3 result’s 2.4x is the most likely misreading of this release.
- Cache-warm runs. If your files sit in DuckDB’s external file cache, or your pipeline re-reads the same objects, there is no remote latency to hide. You would be benchmarking the cache, not S3.
- Wide scans over few, large reads. Reading all four columns of a file with a handful of large row groups approaches sequential bulk transfer, where bandwidth, not request latency, is the bottleneck. Overlapping a few large requests cannot overlap much.
Cross-engine evidence supports the general pattern that engine optimizations land very unevenly. A SQL Server study of bitvector-aware optimization reported workload-level CPU reductions of 22% to 64% but up to two orders of magnitude for individual queries: same engine, same optimization, wildly different per-query outcomes. That is an analogy from a different system, not a DuckDB result, but it is a useful prior against quoting any single multiplier.
A routing summary, with the confidence level made explicit:
| Your workload shape | Expected 2.0 async-I/O benefit | Basis |
|---|---|---|
| Narrow projection, many row groups, remote S3, cache cold | Large (18.8 s to 7.7 s; the 23-file, 13.6 GB read went 11.8 s to 3.9 s) | Measured, author-reported, single machine |
| One large CSV | Roughly 2x (116 s to 55 s) | Measured, author-reported, single machine |
| Lakes of many tiny ~1 MB Parquet files | Near-null (3.7 s vs 3.3 s) | Measured, author-reported, single machine |
| Narrow projection but cache-warm or local files | Small | Inference from mechanism |
| Full-row scans, few large row groups | Small to none | Inference from mechanism |
| CPU-bound (heavy joins, recursion, UDFs) | None from async I/O; separate 2.0 changes may help | Inference; CPU-side gains are a different feature |
| Latency-sensitive production pipeline | Do not run a dev build regardless | DuckDB install page |
A re-benchmark before you quote anything
MotherDuck’s own methodology doubles as a checklist, and it is simple enough to run this week:
- Install 2.0.0-dev beside your current stable, not over it.
- Disable the external file cache in both versions:
SET enable_external_file_cache = false;so every run actually hits S3 (MotherDuck’s post uses exactly this setting, with the comment “so every run really hits S3”). Without it, repeat runs measure the cache and the comparison is meaningless. - Run your real queries, not the Stack Overflow demo. Pick the three to five scans that dominate your pipeline’s S3 time.
- Vary the shape deliberately. Re-run with a wider projection (more columns) and, if you control the files, against a copy rewritten with fewer, larger row groups. The delta between those runs tells you how much of your gain is shape-dependent, which is the thing you actually need to know before committing.
- Repeat each run several times. A single home-connection run has obvious variance; the author flags the numbers as single-machine for a reason.
If your headline queries show a large, stable win on 2.0.0-dev, you have evidence for planning a migration at the stable release. If they show little or nothing, you have saved yourself a migration motivated by someone else’s row-group layout.
Migration risk is the other half of the decision
Even a confirmed 2.4x on your own data does not make the alpha production-safe, and the install page says so outright. Beyond release-channel risk, the preview post lists the concrete breakage surface: a new default storage format (existing DuckDB database files need re-testing, and you should confirm read/write compatibility for any persisted files before upgrading anything that writes them), a reworked C API (any client embedding DuckDB through the C interface should expect code changes), a new SQL parser (edge-case SQL that parsed before may parse differently), and “a small number of carefully chosen breaking changes.” None of these are reasons to avoid 2.0; all of them are reasons to schedule the migration as work, after the stable release, with a rollback path.
One clarification worth making: this analysis concerns DuckDB reading Parquet files from S3, which does not touch the native storage format at all. If your pipeline only ever queries Parquet in place and never writes DuckDB database files, the storage-format change is mostly irrelevant to you, and the parser and client-API changes carry the migration risk. If you do persist DuckDB files, the calculus flips. The MotherDuck result, incidentally, slightly strengthens the case for querying Parquet in place rather than loading it into a warehouse first: if remote scans get cheaper, the “just load it” argument gets weaker. But that economics argument only holds for workload shapes that actually capture the I/O win, which returns everything to the re-benchmark.
Verdict
Route by workload shape, not by a multiplier. The 18.8s-to-7.7s figure is one measured query on one laptop against an alpha, compared against 1.5.5 rather than today’s 1.5.6 stable, from a vendor with a commercial interest in DuckDB momentum, and the author tells you to re-run it yourself. The shapes most likely to benefit are narrow column projections over many small row groups on remote, cache-cold storage, in one file or across many. The shapes least likely to benefit are lakes of many tiny files, the one measured near-null case, along with cache-warm runs, wide bulk scans, and CPU-bound queries; apart from the tiny-files row, that half of the routing table is mechanism-based inference. My recommendation: keep production on 1.5.6 or 1.4.5 LTS, run the cache-disabled benchmark on 2.0.0-dev against your real queries this month, and plan the actual migration for the stable fall release after re-testing storage-format, C-API, and parser breakage. If a lost afternoon of benchmarking is the price of knowing whether the 2.4x applies to your data, that is a cheap price; quoting the figure without running it is not.

Join the discussion
Share a useful perspective or ask a question about this article.