
Databricks ML in Action Course
Master end-to-end machine learning on the Databricks Lakehouse Platform — from data ingestion and feature engineering to distributed training, model serving, and production MLOps. This hands-on course equips data scientists and ML engineers with the tools, workflows, and governance practices that power real enterprise AI systems.
What you will learn:
Configure Databricks clusters, storage integrations, and Unity Catalog for secure ML workflows.
Build scalable Delta Lake pipelines that deliver clean, versioned data ready for model training.
Design and govern reusable feature tables using Databricks Feature Store to eliminate training-serving skew.
Track, register, and promote models through a governed MLflow lifecycle from experiment to production.
Deploy real-time and batch inference endpoints with custom PyFunc logic and automated monitoring.
Implement CI/CD pipelines, drift detection, and automated retraining to operate production ML systems reliably.
How you study in practice Databricks ML in Action Course
How you practise Databricks ML in Action Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsDatabricks Platform Foundations
Databricks Platform Foundations
Lesson 1 • Notebooks and Collaborative Development
Introduces notebook creation, magic commands, and co-authoring features. Establishes reproducible development habits used in all later ML experiments.
Lesson 2 • Workspace Architecture and Navigation
Covers the Databricks control plane, data plane, and workspace UI. Grounds students in the environment they will use throughout every subsequent chapter.
Lesson 3 • Data Access and Storage Integration
Explains mount points, Unity Catalog, and credential management for cloud storage. Enables secure, consistent data access required for ML pipelines.
Lesson 4 • Databricks CLI and REST API Basics
Introduces programmatic workspace control via CLI and REST endpoints. Prepares students to automate tasks and integrate Databricks into external toolchains.
Lesson 5 • Cluster Configuration and Management
Teaches cluster types, runtime versions, and autoscaling policies. Students configure job and interactive clusters suited to ML workloads.
Chapter 2HideHide detailsSee detailsData Engineering for ML Pipelines
Data Engineering for ML Pipelines
Lesson 1 • Delta Live Tables for Pipeline Orchestration
Introduces declarative pipeline definitions, expectations, and lineage tracking. Automates multi-stage data quality enforcement before features reach model training.
Lesson 2 • Data Transformation with Spark SQL
Covers aggregations, window functions, and complex type handling in Spark SQL. Transforms raw data into structured features consumed by ML models.
Lesson 3 • Apache Spark Core Concepts
Covers RDDs, DataFrames, and the Spark execution model. Provides the computational foundation for all data processing tasks in later ML chapters.
Lesson 4 • Delta Lake for ML Workloads
Introduces ACID transactions, time travel, and Delta table optimisation. Ensures data reliability and reproducibility across iterative ML experiments.
Lesson 5 • Data Ingestion Patterns
Teaches batch and streaming ingestion from files, databases, and event streams. Students build reliable source-to-landing-zone pipelines for ML feature creation.
Chapter 3HideHide detailsSee detailsFeature Engineering and the Feature Store
Feature Engineering and the Feature Store
Lesson 1 • Eliminating Training-Serving Skew
Addresses consistency between offline training data and online serving features. Students implement lookup-based training sets that mirror production inference paths.
Lesson 2 • Databricks Feature Store Architecture
Explains Feature Store tables, metadata, and the offline-online split. Students understand how features are stored, versioned, and retrieved for training and serving.
Lesson 3 • Writing and Reading Feature Tables
Teaches the write and read APIs for batch and streaming feature updates. Connects ingestion pipelines from Chapter 2 to downstream model training workflows.
Lesson 4 • Feature Engineering Fundamentals
Covers encoding, scaling, imputation, and interaction terms using Spark. Establishes the feature vocabulary applied throughout the Feature Store and model training chapters.
Lesson 5 • Feature Governance and Reuse
Covers feature discovery, access control, and cross-team sharing via Unity Catalog. Reduces duplicate engineering effort and enforces data lineage across ML projects.
Chapter 4HideHide detailsSee detailsModel Training with MLflow
Model Training with MLflow
Lesson 1 • Training Deep Learning Models
Introduces TensorFlow and PyTorch training loops integrated with MLflow logging. Prepares students for distributed deep learning covered in the advanced training chapter.
Lesson 2 • Hyperparameter Tuning with Hyperopt
Teaches Bayesian optimisation and parallel search using Hyperopt and SparkTrials. Students systematically improve model performance beyond manual grid search.
Lesson 3 • MLflow Tracking Fundamentals
Introduces experiments, runs, parameters, metrics, and tags in MLflow. Establishes the logging discipline that underpins reproducible ML throughout the course.
Lesson 4 • Training Scikit-learn and Spark ML Models
Covers pipeline construction, cross-validation, and MLflow autologging for both frameworks. Students train and log baseline models that serve as benchmarks for advanced techniques.
Lesson 5 • MLflow Projects and Reproducibility
Covers MLflow Projects, entry points, and environment specifications for portable training. Enables teams to reproduce any experiment run on any compatible cluster.
Chapter 5HideHide detailsSee detailsModel Registry and Lifecycle Management
Model Registry and Lifecycle Management
Lesson 1 • Registering and Versioning Models
Teaches programmatic and UI-based model registration from MLflow runs. Students establish a consistent versioning discipline for all trained models.
Lesson 2 • Stage Transitions and Approval Workflows
Covers Staging, Production, and Archived transitions with webhook-based notifications. Implements a controlled promotion process that mirrors enterprise MLOps standards.
Lesson 3 • Unity Catalog Model Registry
Explains the Unity Catalog-backed Model Registry for cross-workspace governance. Students migrate from workspace-level to catalog-level model management for enterprise scale.
Lesson 4 • Model Signatures and Input Examples
Introduces MLflow model signatures, input examples, and schema enforcement. Prevents runtime type errors and documents expected model interfaces for downstream consumers.
Lesson 5 • MLflow Model Registry Overview
Explains registered models, versions, stages, and aliases in the Model Registry. Connects experiment runs from Chapter 4 to a governed promotion pipeline.
Chapter 6HideHide detailsSee detailsScalable and Distributed Model Training
Scalable and Distributed Model Training
Lesson 1 • Distributed Hyperparameter Tuning at Scale
Extends Hyperopt SparkTrials and introduces Optuna for large-scale parallel search. Students tune deep learning models across hundreds of concurrent trials efficiently.
Lesson 2 • TorchDistributor for PyTorch Training
Introduces TorchDistributor as the native Databricks API for distributed PyTorch. Students launch multi-GPU and multi-node PyTorch jobs without manual cluster configuration.
Lesson 3 • Distributed Training Concepts
Covers data parallelism, model parallelism, and gradient synchronization strategies. Provides the theoretical grounding needed before implementing distributed frameworks.
Lesson 4 • Horovod for Distributed Deep Learning
Teaches Horovod initialization, rank-based data sharding, and HorovodRunner on Databricks. Students convert single-node training scripts to multi-node distributed jobs.
Lesson 5 • Pandas UDFs for Distributed Inference
Covers scalar, grouped map, and iterator Pandas UDFs for applying models at Spark scale. Enables batch inference on billions of rows without moving data off the cluster.
Chapter 7HideHide detailsSee detailsModel Deployment and Real-Time Serving
Model Deployment and Real-Time Serving
Lesson 1 • Databricks Model Serving Architecture
Explains serverless model serving, endpoint types, and traffic routing. Orients students to the serving layer before configuring and testing live endpoints.
Lesson 2 • External Deployment Targets
Covers exporting models to ONNX, Docker containers, and cloud-native serving platforms. Prepares students to deploy Databricks-trained models outside the Databricks environment.
Lesson 3 • Custom Inference Logic with PyFunc
Teaches the MLflow PyFunc flavor for wrapping arbitrary Python logic as a deployable model. Students add preprocessing, postprocessing, and ensemble logic to serving endpoints.
Lesson 4 • Deploying MLflow Models to Endpoints
Covers endpoint creation, model version pinning, and REST API invocation. Students deploy registered models from Chapter 5 and validate responses end-to-end.
Lesson 5 • Batch Inference Pipelines
Implements scheduled batch scoring using Databricks Jobs and mlflow.pyfunc.spark_udf. Complements real-time serving for high-throughput, latency-tolerant use cases.
Chapter 8HideHide detailsSee detailsMLOps, Monitoring, and Production Governance
MLOps, Monitoring, and Production Governance
Lesson 1 • Audit, Lineage, and Compliance
Covers Unity Catalog lineage, MLflow audit logs, and access governance for regulated environments. Students document model provenance to satisfy internal and external audit requirements.
Lesson 2 • Lakehouse Monitoring for ML
Introduces Databricks Lakehouse Monitoring for automated metric computation on inference tables. Provides out-of-the-box dashboards for model quality and data health.
Lesson 3 • MLOps Principles and Maturity Levels
Defines MLOps maturity from manual experimentation to fully automated pipelines. Frames the governance and automation goals students implement in subsequent sections.
Lesson 4 • Data and Model Drift Detection
Covers statistical tests for covariate shift, label drift, and prediction drift. Students implement drift alerts that trigger retraining pipelines automatically.
Lesson 5 • Automated Retraining Pipelines
Builds trigger-based and scheduled retraining workflows using Databricks Jobs and Delta Live Tables. Keeps production models current without manual intervention.
Your valid completion certificate
This course is for you:
Data scientists ready to move beyond local notebooks into cloud-scale ML.
ML engineers seeking structured workflows for deploying and monitoring models.
Software engineers transitioning into machine learning roles on cloud platforms.
Analytics engineers who want to extend their data pipelines into model training.
Graduate students applying academic ML knowledge to real enterprise environments.
BI developers curious about building predictive systems on top of their data.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















