This repository contains a complete, real-world Data Lakehouse implementation built on Databricks, including datasets, notebooks, SQL examples. Everything here is designed to help you understand how modern data teams use Databricks in practice, from data ingestion and transformation to analytics-ready data products.
This project follows the Medallion Architecture:
- Raw data ingestion
- Schema inference and storage as Delta tables
- Data cleaning and standardization
- Type casting and validation
- Dimensional Data Model (Business Transformation)
- Ready for BI and analysis
- Databricks
- Apache Spark
- PySpark
- Spark SQL
- Delta Lake
- Unity Catalog
- Basic SQL, Python and some Pyspark knowledge
- No prior Databricks experience required
This project is licensed under the MIT License. You are free to use, modify, and share this project with proper attribution.