For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's alright, it's pretty alpha software. On v1.5.4, catalog filtered counts are broken, afaik. I went to main/v2 to fix it, and then the SQL parser in duckdb v2 is 10x slower, which was another wrench in the gears. It's been a bit of a pain tbh
It was not popular then (https://www.tiobe.com/tiobe-index/), and it was not selected by DuckDB. IF they feel C++ is holding them back they can rewrite it...
Don't think it's widely known or understood, but Ducklake doesn't require duckdb: it's just a (good) data lake spec that works in duckdb.
There's a cool alternate rust/datafusion ecosystem initiative going on at https://github.com/datafusion-contrib/datafusion-ducklake, and think the Quack protocol opens up a lot of cool possibilities too.
If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
Thanks for the shoutout.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
Motherduck is offering a free copy of O'reilly's "DuckLake: The Definitive Guide" book on their DuckLake web page
https://motherduck.com/product/ducklake/
Is this basically a table format like Delta/Iceberg but with an SQL engine built in via DuckDB?
It's alright, it's pretty alpha software. On v1.5.4, catalog filtered counts are broken, afaik. I went to main/v2 to fix it, and then the SQL parser in duckdb v2 is 10x slower, which was another wrench in the gears. It's been a bit of a pain tbh
Yep, they made the spec 1.0 but it isn’t 1.0 software. Browse the bugs before use.
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
Absolutely. I love it, but you need a fork for now.
But "Duckpond" was right there!
DuckDB is the best thing world got since 1990s.
https://duckdb.org/2025/05/19/the-lost-decade-of-small-data....
Why program greenfield in C++? I know "made with rust" is a meme but seriously, why not a memory safe language in 2026?
DuckDB has been around since 2018. Naturally its offspring use C++.
Rust has been out since 2010. It was recognized and popular by 2017.
It was not popular then (https://www.tiobe.com/tiobe-index/), and it was not selected by DuckDB. IF they feel C++ is holding them back they can rewrite it...
[flagged]