Your Data Lakehouse Magazine and Community
News, newsletters, and a thriving community for everyone building on the data lakehouse ecosystem. Stay informed with weekly roundups, deep dives into Apache Iceberg, and expert guides from practitioners.
- 400+
- Articles & tutorials
- 300+
- Glossary entries
- 5
- Free lakehouse books
- Weekly
- Newsletter & roundups
Find your way around
Four ways into the Hub — reference, reading, watching, and meeting people.
Knowledge Base
A searchable glossary of lakehouse, Iceberg, and catalog terminology.
BrowseThe Blog
Tutorials, architecture deep dives, and ecosystem news, published weekly.
BrowseVideo Explainers
Short, focused walkthroughs of the concepts that are hard to read about.
BrowseEvents Calendar
Meetups, webinars, and Lakehouse Linkups happening across the community.
BrowseLatest Posts
Designing Batch Pipelines That Write Well Into Apache Iceberg
How to design batch pipelines that write well into Apache Iceberg: commit strategy, partitioning, sort order, write-audit-publish, and maintenance done right.
Read more
Apache Iceberg Support Across the Major Hyperscalers
How AWS, Google Cloud, and Microsoft Azure actually support Apache Iceberg: storage, catalogs, maintenance, governance, and interoperability, layer by layer.
Read more
Apache Polaris 1.7.0 and the Quiet Work of Making a Catalog Trustworthy
Apache Polaris 1.7.0 deep dive: idempotent writes, semantic models, stricter credential vending, orphan cleanup, and what the upgrade asks of you.
Read moreMust Read Articles
Deep dives into Apache Iceberg, agentic AI, and modern data lakehouse architecture by Alex Merced.
What is Apache Iceberg? The Table Format Revolution
Learn how Apache Iceberg turned raw Parquet files in S3 into a fully ACID-compliant, time-traveling analytical database without moving your data out of object storage.
Read articleAgentic Analytics on the Apache Lakehouse
How autonomous AI agents replace manual dashboard querying by reading governed semantic layers directly on your data lakehouse.
Read article EcosystemThe 2025 State of the Apache Iceberg Ecosystem
Survey data from data professionals on adoption rates, popular tooling, and where the ecosystem is heading through 2026.
Read article Data LakehouseWhat Are Table Formats and Why Were They Needed?
Table formats solved the ACID, schema evolution, and query performance problems that turned data lakes into unmanageable swamps.
Read article Apache Iceberg2026 Intro to Apache Iceberg
A beginner-to-intermediate introduction covering the metadata layer, catalog integrations, and why Iceberg became the dominant table format.
Read article Agentic AIWhy AI Fails Without a Semantic Layer
The business context and governed metrics AI agents need to generate accurate, trustworthy analytical answers at scale.
Read articleWhy Build a Data Lakehouse?
Unified Access
Eliminate data silos. Query your data where it lives—in S3, ADLS, or GCS—without moving it.
Open Standards
Avoid vendor lock-in. Use open formats like Apache Iceberg and open catalogs to keep your data accessible to any engine.
High Performance
Achieve sub-second query performance on data lake scale datasets using engines like Dremio.
The Modern Data Stack is Open
The Data Lakehouse Hub is your central resource for tutorials, architectural guides, and community support. Whether you're migrating from a warehouse or building from scratch, we have the resources to help you succeed with open data standards.
Meet the author
Data Lakehouse Reading
Full-length books on Apache Iceberg, Apache Polaris, and lakehouse architecture — no paywall.
Stay Ahead of the Curve
Subscribe to our newsletter and event calendar to get the latest tutorials, webinars, and meetups delivered to your inbox.