
Databricks training
Master the Databricks platform from the ground up — covering the lakehouse architecture, Apache Spark, Delta Lake, and ML workflows. This training equips data engineers, analysts, and data scientists with the hands-on skills needed to build production-grade pipelines and analytics solutions. If you work with data at scale, this is the training that moves your career forward.
What you will learn:
This course covers every major layer of the Databricks platform, starting with the lakehouse architecture and workspace fundamentals. You will write and optimise Apache Spark jobs, manage Delta Lake tables, and build reliable ETL pipelines using Auto Loader and Delta Live Tables. You will apply the medallion architecture to organise data into Bronze, Silver, and Gold layers. The course also covers Databricks SQL for analytics, MLflow for machine learning lifecycle management, and workflow orchestration for production deployments. Advanced topics include Spark performance tuning, Unity Catalog governance, real-time Kafka streaming, and generative AI with Mosaic AI.
How you study in practice Databricks training
How you practise Databricks training
For companies looking to train their teams
With Dedika for Businesses, the course includes exercises and examples tailored to your own business and the specific needs of your company.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsIntroduction to Databricks and the Lakehouse
Introduction to Databricks and the Lakehouse
Lesson 1 • Navigating the Databricks Workspace
Covers the workspace UI, resource organization, and user roles. Learners gain hands-on familiarity with the environment before writing any code.
Lesson 2 • Connecting to Data Sources
Demonstrates mounting cloud storage and connecting to external databases. Establishes data access patterns required for all data engineering and analytics tasks.
Lesson 3 • What Is the Lakehouse Architecture
Defines the lakehouse paradigm and contrasts it with data warehouses and data lakes. Establishes the architectural context for all subsequent Databricks work.
Lesson 4 • Working with Databricks Notebooks
Introduces notebook creation, magic commands, and multi-language support. Notebooks are the primary development interface used in every subsequent chapter.
Lesson 5 • Clusters: Configuration and Management
Explains cluster types, configuration options, and lifecycle management. Proper cluster setup is essential for efficient resource use throughout the course.
Chapter 2HideHide detailsSee detailsApache Spark Fundamentals in Databricks
Apache Spark Fundamentals in Databricks
Lesson 1 • Reading and Writing Data Formats
Covers reading and writing Parquet, JSON, CSV, and Delta formats with options. Proper format handling is foundational for all ingestion and output pipelines.
Lesson 2 • DataFrames and the DataFrame API
Covers DataFrame creation, schema inference, and core transformations. DataFrames are the primary abstraction used throughout all data processing chapters.
Lesson 3 • Spark Performance Basics
Introduces partitioning, caching, and broadcast joins for performance tuning. These techniques are applied in advanced optimization chapters later in the course.
Lesson 4 • Spark Architecture and Execution Model
Explains drivers, executors, DAGs, and lazy evaluation. Understanding execution flow is critical for writing efficient Spark code in later chapters.
Lesson 5 • Spark SQL for Data Querying
Teaches SQL-based querying on top of Spark using temporary views and catalog tables. Bridges SQL skills with the Spark execution engine.
Chapter 3HideHide detailsSee detailsDelta Lake: Reliable Data Storage
Delta Lake: Reliable Data Storage
Lesson 1 • DML Operations on Delta Tables
Teaches INSERT, UPDATE, DELETE, and MERGE operations on Delta tables. Enables learners to implement upsert patterns essential for incremental data pipelines.
Lesson 2 • Optimizing Delta Tables
Covers OPTIMIZE, ZORDER, and VACUUM commands for performance and storage management. These operations are critical for maintaining production-grade Delta tables.
Lesson 3 • Time Travel and Data Versioning
Demonstrates querying historical table versions and restoring prior states. Time travel supports auditing and error recovery in production data pipelines.
Lesson 4 • Creating and Managing Delta Tables
Covers DDL operations, table properties, and managed vs. external tables. These skills are required for building structured data layers in later pipeline chapters.
Lesson 5 • Delta Lake Core Concepts
Explains the Delta transaction log, ACID guarantees, and how Delta differs from plain Parquet. Provides the conceptual foundation for all Delta table operations.
Chapter 4HideHide detailsSee detailsData Ingestion and ETL Pipelines
Data Ingestion and ETL Pipelines
Lesson 1 • Auto Loader for Incremental Ingestion
Introduces Auto Loader's file notification and directory listing modes for scalable ingestion. Auto Loader is the recommended Databricks pattern for continuous file-based ingestion.
Lesson 2 • Batch Ingestion Patterns
Covers full and incremental batch loading strategies using Spark and Delta. Establishes the ingestion patterns that underpin all pipeline architectures in this chapter.
Lesson 3 • Building Transformation Logic
Covers multi-step transformation patterns including cleansing, enrichment, and deduplication. Transformation logic is the core value-add layer in any ETL pipeline.
Lesson 4 • Delta Live Tables for Pipeline Orchestration
Introduces Delta Live Tables (DLT) for declarative pipeline definition and quality enforcement. DLT simplifies pipeline management and integrates with the medallion architecture.
Lesson 5 • Structured Streaming Fundamentals
Teaches micro-batch and continuous streaming with Spark Structured Streaming. Provides the streaming foundation needed for real-time pipeline sections ahead.
Chapter 5HideHide detailsSee detailsData Modeling and the Medallion Architecture
Data Modeling and the Medallion Architecture
Lesson 1 • Medallion Architecture Principles
Defines Bronze, Silver, and Gold layer responsibilities and data contracts. Provides the structural framework applied in every subsequent modeling and pipeline section.
Lesson 2 • Designing the Bronze Layer
Covers raw data landing, schema-on-read, and preserving source fidelity. A well-designed Bronze layer ensures full auditability and reprocessing capability.
Lesson 3 • Constructing Gold Layer Aggregates
Covers building pre-aggregated, denormalized tables optimized for BI and reporting. Gold tables reduce query complexity and improve dashboard performance.
Lesson 4 • Building the Silver Layer
Teaches cleansing, type casting, deduplication, and conforming data to a unified schema. The Silver layer is the primary source for analytical and ML workloads.
Lesson 5 • Data Quality and Expectations
Introduces constraint enforcement, quarantine patterns, and quality metrics tracking. Embedding quality checks at each layer prevents bad data from propagating downstream.
Chapter 6HideHide detailsSee detailsDatabricks SQL and Analytics
Databricks SQL and Analytics
Lesson 1 • Query Optimization for Analytics
Teaches query plan analysis, caching strategies, and result caching for SQL workloads. Optimized queries reduce warehouse costs and improve user experience.
Lesson 2 • Building Dashboards and Visualizations
Covers creating charts, dashboards, and alerts within Databricks SQL. Dashboards enable business stakeholders to consume data without writing code.
Lesson 3 • Unity Catalog for Data Governance
Introduces Unity Catalog's three-level namespace, access control, and data lineage. Governance is foundational for secure, compliant analytics across teams.
Lesson 4 • Advanced SQL Techniques in Databricks
Teaches window functions, CTEs, and higher-order functions for complex analytical queries. These techniques enable sophisticated analysis directly within Databricks SQL.
Lesson 5 • SQL Warehouses and Query Execution
Covers SQL warehouse types, sizing, and query routing. Proper warehouse configuration directly impacts query performance and cost for analytics workloads.
Chapter 7HideHide detailsSee detailsMachine Learning with Databricks and MLflow
Machine Learning with Databricks and MLflow
Lesson 1 • MLflow Model Registry
Introduces model registration, stage transitions, and version management in the Model Registry. The registry provides governance and traceability for production models.
Lesson 2 • Model Serving and Batch Inference
Covers real-time model serving endpoints and batch inference with Spark. Deploying models in both modes ensures flexibility for diverse business requirements.
Lesson 3 • Training Models at Scale with Spark
Covers distributed training using pandas UDFs, Spark ML, and Hyperopt for tuning. Distributed training reduces time-to-insight on large datasets.
Lesson 4 • Experiment Tracking with MLflow
Teaches logging parameters, metrics, and artifacts with MLflow Tracking. Systematic experiment tracking enables reproducibility and informed model selection.
Lesson 5 • Databricks ML Runtime and Feature Engineering
Covers the ML Runtime environment, pre-installed libraries, and Feature Store basics. A well-prepared feature set is the foundation of reproducible model training.
Chapter 8HideHide detailsSee detailsWorkflow Orchestration and Production Deployment
Workflow Orchestration and Production Deployment
Lesson 1 • Monitoring, Logging, and Cost Management
Covers cluster event logs, job metrics, audit logs, and DBU cost tracking. Observability and cost control are essential for sustainable production operations.
Lesson 2 • Scheduling and Triggering Workflows
Teaches cron-based scheduling, file-arrival triggers, and external API triggers. Flexible triggering ensures pipelines run at the right time with minimal manual intervention.
Lesson 3 • Infrastructure as Code with Terraform
Introduces provisioning Databricks resources using the Terraform provider. Infrastructure as code ensures reproducible, auditable environment setup across workspaces.
Lesson 4 • Databricks Jobs and Task Orchestration
Covers creating multi-task jobs, setting dependencies, and configuring compute for each task. Jobs are the primary mechanism for running production workloads on a schedule.
Lesson 5 • CI/CD for Databricks with Repos
Covers Git integration via Databricks Repos, branch strategies, and automated deployment. CI/CD practices reduce deployment risk and accelerate delivery of data products.
Your valid completion certificate
This course is for you:
Data Engineer: ready to move beyond basic pipelines into a unified lakehouse platform.
Data Analyst: wants to query and visualise data at scale using Databricks SQL.
Data Scientist: needs a reliable platform for feature engineering and model deployment.
Analytics Engineer: building dbt-style transformation layers and exploring Spark-native alternatives.
Cloud Engineer: supporting data teams and needing deeper Databricks infrastructure knowledge.
Career Changer: transitioning from software development into the modern data engineering field.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.

Top trainings
FAQs
Who is Dedika?
Is the certificate valid in Pakistan?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















