Query Huge Datasets With a Tiny Database File

TL;DR: A DuckDB database file can act as a portable catalog for remote data without storing any of it. This lets teams share and query massive datasets stored elsewhere using just a tiny, shareable file.
Key facts
- Category
- Database
- Impact
- High
- Published
- Source
- DuckDB Blog
Full summary
A new pattern uses tiny, data-less DuckDB files as shareable catalogs to query massive datasets stored remotely in object storage.
A new architectural pattern is changing how developers think about databases, using a file that contains no data at all. According to a post on the DuckDB Blog, which credits the idea to developer Nikolas Goebel, a DuckDB database file can function purely as a portable data catalog. Instead of storing tables and rows, the file holds only a collection of view definitions. These views point to external data files, such as Parquet files stored in cloud object storage like Amazon S3. The result is a remarkably lightweight file, often just a few hundred kilobytes, that acts as a complete, queryable map to potentially petabytes of remote data. Anyone who receives this tiny file can attach it to their local DuckDB instance and immediately begin running complex SQL queries against massive datasets as if they were all stored on their own machine.
AThe mechanism behind this pattern is elegant in its simplicity, leveraging a standard SQL feature: the `VIEW`. A view is essentially a stored query that the database presents as a virtual table. In this data-less approach, the DuckDB file is populated with a series of `CREATE VIEW` statements. Each statement defines a table schema and, crucially, specifies the location of the underlying data using a path to a remote file or a set of files. For example, a view named `sales_q1` might be defined by a query like `SELECT * FROM 's3://my-company-data/sales/q1/*.parquet'`. When a user runs a query against this `sales_q1` view, DuckDB’s powerful query engine intercepts the request. It understands that the data resides remotely, so it transparently connects to the object storage, streams only the necessary data for the query, and processes it on the fly. The database file itself never holds the data; it is merely a collection of recipes for finding and structuring it.
This approach is a powerful real-world application of a core principle in modern data systems: the separation of compute and storage. Traditional database systems often tightly couple the engine that processes queries (compute) with the disk system that holds the data (storage). By using a data-less file, this pattern introduces a third, decoupled component: the catalog. The catalog (the DuckDB file) is separate from the storage (the Parquet files in the cloud) and the compute (the user's local machine running DuckDB). This modularity is a hallmark of the modern data stack, offering greater flexibility, scalability, and cost-efficiency. It provides a serverless, zero-infrastructure alternative to heavier, centrally managed catalog systems like AWS Glue or Hive Metastore, which require dedicated services to run. For teams that need a simple, shareable way to define a common view over a data lake, this pattern is a game-changer.
The practical implications for data teams are significant, primarily in improving collaboration and agility. A data engineering team can now curate a definitive set of tables from a chaotic data lake, encode them as views in a single DuckDB file, and commit that file to a Git repository. This makes the data catalog version-controlled, reviewable, and easily distributable. Analysts and data scientists can simply pull the latest version of the file and start their work immediately, without needing to manually track file paths or manage credentials for dozens of different data sources. This democratizes data access and dramatically lowers the barrier to performing ad-hoc exploration and analysis. As this pattern gains adoption, the next steps will involve developing best practices for managing schema evolution, handling access control for sensitive data, and integrating these portable catalogs into larger data governance frameworks.
Why it matters
This pattern simplifies data access in modern data stacks. It provides a lightweight, serverless way to create a unified view over distributed data sources, decoupling the data catalog from compute and storage and making collaboration easier for data teams.
Business impact
This approach can reduce data management overhead and infrastructure costs. By using a small, portable file as a data catalog, companies can avoid running dedicated catalog services and enable faster, more flexible analysis without moving large datasets.
Tags
Related on Notifire
Related stories
Primary source: DuckDB Blog