
Building Modern Data Applications Using Databricks Lakehouse Course
Master the Databricks Lakehouse platform from architecture fundamentals to production-grade pipelines. Learn to ingest, transform, and govern data at scale using Delta Lake, Unity Catalog, and MLflow. This course equips data engineers and architects with the hands-on skills to build reliable, cost-efficient modern data applications.
What you will learn:
Configure Databricks clusters, notebooks, and cloud storage integrations for production workloads.
Build scalable ingestion pipelines using Auto Loader, COPY INTO, and Structured Streaming.
Implement Medallion Architecture with Bronze, Silver, and Gold Delta Lake layers.
Apply Unity Catalog policies to enforce data governance, auditing, and secure data sharing.
Optimise Spark SQL queries and Delta tables using Z-ORDER, partitioning, and caching strategies.
Integrate MLflow tracking and the Model Registry to manage the full machine learning lifecycle.
How you study practically Building Modern Data Applications Using Databricks Lakehouse Course
How you practise Building Modern Data Applications Using Databricks Lakehouse Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsLakehouse Architecture and Databricks Fundamentals
Lakehouse Architecture and Databricks Fundamentals
Lesson 1 • Databricks Platform Overview
Maps the major Databricks services: workspace, clusters, jobs, and Unity Catalog. Students gain orientation before configuring any resource.
Lesson 2 • Cluster Configuration and Management
Covers all-purpose vs. job clusters, autoscaling, and runtime selection. Correct cluster design directly reduces cost and improves pipeline reliability.
Lesson 3 • Notebooks and Collaborative Development
Introduces multi-language notebooks, magic commands, and co-authoring features. Establishes the primary development environment used throughout the course.
Lesson 4 • Storage Integration and Cloud Connectivity
Explains how Databricks mounts and accesses cloud object storage securely. Students configure storage credentials needed for all subsequent data exercises.
Lesson 5 • Evolution from Data Lakes to Lakehouse
Traces the limitations of data warehouses and data lakes that motivated the Lakehouse design. Provides the historical context needed to justify every architectural decision covered later.
Chapter 2HideHide detailsSee detailsDelta Lake Internals and Table Management
Delta Lake Internals and Table Management
Lesson 1 • Table Optimisation Techniques
Covers OPTIMISE, Z-ORDER, and file compaction to improve query performance. These techniques are applied to production tables in advanced chapters.
Lesson 2 • Creating and Modifying Delta Tables
Covers DDL for managed and external tables, schema definition, and ALTER operations. Students build the tables used in all pipeline exercises.
Lesson 3 • DML Operations and Merge Patterns
Teaches INSERT, UPDATE, DELETE, and MERGE INTO with practical upsert scenarios. MERGE is the cornerstone of incremental data loading covered in later chapters.
Lesson 4 • Time Travel and Data Versioning
Demonstrates querying historical table versions and restoring prior states. Students use time travel for auditing and recovery scenarios.
Lesson 5 • Delta Lake Core Concepts
Explains the transaction log, Parquet data files, and how ACID guarantees are achieved. This foundation is required before any table operation is meaningful.
Chapter 3HideHide detailsSee detailsData Ingestion Patterns and Auto Loader
Data Ingestion Patterns and Auto Loader
Lesson 1 • Batch Ingestion with COPY INTO
Introduces COPY INTO for idempotent file loading from cloud storage into Delta tables. Establishes the simplest ingestion pattern before introducing streaming alternatives.
Lesson 2 • Auto Loader Architecture
Explains how Auto Loader detects new files using directory listing or file notification. Students understand the internal mechanics before configuring pipelines.
Lesson 3 • Ingesting Semi-Structured Data
Teaches parsing JSON, XML, and nested structures during ingestion. Students flatten and normalise complex payloads into queryable Delta columns.
Lesson 4 • Streaming Ingestion from Message Brokers
Connects Databricks to event streaming platforms for real-time ingestion into Delta. Students configure offsets, consumer groups, and exactly-once semantics.
Lesson 5 • Schema Inference and Evolution
Covers automatic schema detection, schema hints, and handling schema changes at runtime. Robust schema management prevents pipeline failures in production.
Chapter 4HideHide detailsSee detailsSpark SQL and DataFrame Transformations
Spark SQL and DataFrame Transformations
Lesson 1 • User-Defined Functions and Higher-Order Functions
Introduces Python UDFs, Pandas UDFs, and built-in higher-order functions for arrays and maps. Students extend Spark SQL with custom logic efficiently.
Lesson 2 • DataFrame API in Python and Scala
Teaches DataFrame creation, column operations, and chaining transformations programmatically. Students choose between SQL and API approaches based on use case.
Lesson 3 • Spark SQL Fundamentals
Reviews SQL syntax within Databricks, including CTEs, window functions, and subqueries. Provides the query foundation for all transformation work in the course.
Lesson 4 • Aggregations and Joins
Covers groupBy, pivot, and all join strategies with their performance implications. Correct join selection is critical for avoiding data skew in large datasets.
Lesson 5 • Query Optimisation and the Catalyst Planner
Explains logical and physical plans, predicate pushdown, and partition pruning. Students read explain plans to diagnose and fix slow queries.
Chapter 5HideHide detailsSee detailsBuilding Medallion Architecture Pipelines
Building Medallion Architecture Pipelines
Lesson 1 • Delta Live Tables for Pipeline Orchestration
Introduces DLT declarative pipelines, expectations, and automatic dependency resolution. Students migrate a manual Medallion pipeline to a managed DLT pipeline.
Lesson 2 • Bronze Layer Implementation
Builds the raw ingestion layer using Auto Loader with metadata columns and audit fields. Students apply ingestion patterns from Chapter 3 in a structured pipeline.
Lesson 3 • Medallion Architecture Design Principles
Defines the purpose and data quality expectations of each Medallion layer. Correct layer design prevents data quality issues from propagating downstream.
Lesson 4 • Silver Layer Cleansing and Conforming
Applies data quality rules, deduplication, and type casting to produce Silver tables. Students use MERGE to incrementally update Silver from Bronze.
Lesson 5 • Gold Layer Aggregation and Serving
Creates aggregated, denormalised Gold tables optimised for BI and ML consumption. Students apply Z-ORDER and partitioning for query performance.
Chapter 6HideHide detailsSee detailsUnity Catalog and Data Governance
Unity Catalog and Data Governance
Lesson 1 • Sharing Data with Delta Sharing
Configures open-protocol Delta Sharing to distribute data to external recipients securely. Students publish shares and manage recipient credentials without copying data.
Lesson 2 • Data Lineage and Discovery
Demonstrates automatic column-level lineage capture and the data catalog search interface. Lineage enables impact analysis before schema changes are deployed.
Lesson 3 • Identity and Access Management
Covers users, groups, service principals, and privilege grants in Unity Catalog. Students implement least-privilege access for analysts, engineers, and automated jobs.
Lesson 4 • Unity Catalog Object Model
Explains the metastore, catalog, schema, and table hierarchy in Unity Catalog. Understanding the object model is prerequisite to assigning any permission.
Lesson 5 • Auditing and Compliance Monitoring
Queries the Unity Catalog audit log to detect unauthorised access and policy violations. Students build audit dashboards aligned with data governance requirements.
Chapter 7HideHide detailsSee detailsMachine Learning Integration with MLflow
Machine Learning Integration with MLflow
Lesson 1 • Batch and Real-Time Model Inference
Deploys registered models for batch scoring on Delta tables and real-time serving via endpoints. Students write inference results back to Gold-layer Delta tables.
Lesson 2 • Model Training on Spark DataFrames
Trains scikit-learn and Spark MLlib models using Delta-backed feature data at scale. Students apply cross-validation and hyperparameter tuning within Databricks.
Lesson 3 • MLflow Model Registry and Lifecycle
Registers models, manages stage transitions, and enforces approval workflows in the Model Registry. Students promote a model from staging to production safely.
Lesson 4 • MLflow Tracking and Experiment Management
Teaches logging parameters, metrics, and artefacts to MLflow experiments from Databricks notebooks. Systematic tracking is the foundation of reproducible ML.
Lesson 5 • Feature Engineering with Feature Store
Creates and publishes feature tables in the Databricks Feature Store backed by Delta. Students link features to training datasets for consistent train-serve parity.
Chapter 8HideHide detailsSee detailsProduction Operations and Performance Tuning
Production Operations and Performance Tuning
Lesson 1 • Databricks Workflows and Job Orchestration
Builds multi-task jobs with dependencies, conditional branching, and retry logic in Databricks Workflows. Students replace ad-hoc notebook runs with governed, scheduled pipelines.
Lesson 2 • Caching, Partitioning, and File Management
Applies Delta caching, disk caching, and strategic partitioning to accelerate repeated queries. Students balance partition granularity against small-file overhead.
Lesson 3 • Observability and Alerting
Configures Databricks system tables, Ganglia metrics, and external alerting for pipeline health. Students build a monitoring dashboard covering latency, errors, and data freshness.
Lesson 4 • Cost Management and Cluster Optimisation
Analyses DBU consumption, applies instance pool strategies, and right-sizes clusters. Cost control is essential before scaling workloads to production volumes.
Lesson 5 • Spark Performance Diagnostics
Uses the Spark UI, stage timelines, and task metrics to identify bottlenecks. Students diagnose shuffle spills, skew, and executor failures systematically.
Your valid completion certificate
This course is for you:
Data engineer: ready to move beyond basic Spark into Lakehouse architecture.
Cloud architect: designing unified storage strategies for enterprise data platforms.
Analytics engineer: bridging the gap between raw data and business-ready tables.
Data scientist: wanting to operationalise models within a governed data environment.
BI developer: seeking deeper pipeline ownership beyond dashboards and reports.
Software engineer: transitioning into data infrastructure and platform engineering roles.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.

Top training programmes
FAQ
Who is Dedika?
Is the certificate valid in Kenya?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















