If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
eddietejeda 11 hours ago [-]
Thanks for the shoutout.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
prpl 9 hours ago [-]
What latencies were you targeting?
eddietejeda 6 hours ago [-]
Extremely high concurrency at sub-second response times.
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
paragraft 4 hours ago [-]
That's a tabbed interface on the site that just defaults to postgres. SQLite and duckdb are supported too.
celias 1 days ago [-]
Motherduck is offering a free copy of O'reilly's "DuckLake: The Definitive Guide" book on their DuckLake web page
I think this is an elegant design that's superior to the competition, but I think lakehouses are not as generally useful as vendors would like us to believe. The access controls are limited to what's possible on the underlying bucket.
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
It's alright, it's pretty alpha software. On v1.5.4, catalog filtered counts are broken, afaik. I went to main/v2 to fix it, and then the SQL parser in duckdb v2 is 10x slower, which was another wrench in the gears. It's been a bit of a pain tbh
jauco 12 hours ago [-]
Yep, they made the spec 1.0 but it isn’t 1.0 software. Browse the bugs before use.
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
engineeringwoke 11 hours ago [-]
Absolutely. I love it, but you need a fork for now.
Is this basically a table format like Delta/Iceberg but with an SQL engine built in via DuckDB?
Lucasoato 10 hours ago [-]
Nope, from my understanding the delta log (the files that say which of your data files are actually valid or not) isn’t saved in json/parquet but directly in a database.
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)
snapetom 1 hours ago [-]
Ah, I see. Thanks. Having the metadata in a DB sounds a lot more robust.
loufe 13 hours ago [-]
Why program greenfield in C++? I know "made with rust" is a meme but seriously, why not a memory safe language in 2026?
esafak 12 hours ago [-]
DuckDB has been around since 2018. Naturally its offspring use C++.
OutOfHere 11 hours ago [-]
Rust has been out since 2010. It already was rated "most-loved" by 2016. Systems programmers had adopted it by 2018.
meredithbloom 5 hours ago [-]
And the compiler was horrendously slow for big projects on x86_84 until 2023.
There's a cool alternate rust/datafusion ecosystem initiative going on at https://github.com/datafusion-contrib/datafusion-ducklake, and think the Quack protocol opens up a lot of cool possibilities too.
If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
Here is our write up on the Ducklake blog: https://ducklake.select/2026/07/29/bringing-ducklake-to-data...
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
https://motherduck.com/product/ducklake/
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
https://www.tomwphillips.co.uk/2026/08/the-benefits-of-data-...
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
https://duckdb.org/2025/05/19/the-lost-decade-of-small-data....
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)