HelloStack
DuckLake logo

Integrated data lake and catalog format

Social media
@GitHub User

About DuckLake

DuckLake is a lakehouse format built on SQL that stores all metadata in a SQL database and uses Parquet files for storage. It targets data teams who want lakehouse capabilities without the operational overhead, offering an open, standalone format from the DuckDB team.

DuckLake screenshot

Key Features

DuckLake pairs a SQL catalog with Parquet storage to support lakehouse-like workflows for analysts and engineers:

Catalog database

All metadata lives in a SQL database (PostgreSQL, MySQL, SQLite, or DuckDB) with no custom catalog server required.

Open Parquet storage

Data lives as Parquet files on local disk or object storage, compatible with Iceberg.

Snapshots and time travel

Unlimited snapshots and time-travel queries for historical views without heavy compaction.

ACID transactions

Concurrent access with ACID guarantees across multi-table operations.

Summary

Best for data engineering, analytics, and BI teams that want lakehouse-style features without managing a separate catalog server.

More in Data Pipelines

See all →