Choose your language
Databricks Training
More than 2 million students worldwide

Databricks Training

Master the Databricks platform from the ground up — covering the lakehouse architecture, Apache Spark, Delta Lake, and ML workflows. This training equips data engineers, analysts, and data scientists with the hands-on skills needed to build production-grade pipelines and analytics solutions. If you work with data at scale, this is the training that moves your career forward.

Dedika for businesses

What you will learn:

This course covers every major layer of the Databricks platform, starting with the lakehouse architecture and workspace fundamentals. You will write and optimize Apache Spark jobs, manage Delta Lake tables, and build reliable ETL pipelines using Auto Loader and Delta Live Tables. You will apply the medallion architecture to organize data into Bronze, Silver, and Gold layers. The course also covers Databricks SQL for analytics, MLflow for machine learning lifecycle management, and workflow orchestration for production deployments. Advanced topics include Spark performance tuning, Unity Catalog governance, real-time Kafka streaming, and generative AI with Mosaic AI.

How you study in practice Databricks Training

How you practice Databricks Training

For companies that want to train their team

With Dedika for Business, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Introduction to Databricks and the Lakehouse

  • Lesson 1 • Navigating the Databricks Workspace

    Covers the workspace UI, resource organization, and user roles. Learners gain hands-on familiarity with the environment before writing any code.

  • Lesson 2 • Connecting to Data Sources

    Demonstrates mounting cloud storage and connecting to external databases. Establishes data access patterns required for all data engineering and analytics tasks.

  • Lesson 3 • What Is the Lakehouse Architecture

    Defines the lakehouse paradigm and contrasts it with data warehouses and data lakes. Establishes the architectural context for all subsequent Databricks work.

  • Lesson 4 • Working with Databricks Notebooks

    Introduces notebook creation, magic commands, and multi-language support. Notebooks are the primary development interface used in every subsequent chapter.

  • Lesson 5 • Clusters: Configuration and Management

    Explains cluster types, configuration options, and lifecycle management. Proper cluster setup is essential for efficient resource use throughout the course.

Chapter 2See details

Apache Spark Fundamentals in Databricks

  • Lesson 1 • Reading and Writing Data Formats

    Covers reading and writing Parquet, JSON, CSV, and Delta formats with options. Proper format handling is foundational for all ingestion and output pipelines.

  • Lesson 2 • DataFrames and the DataFrame API

    Covers DataFrame creation, schema inference, and core transformations. DataFrames are the primary abstraction used throughout all data processing chapters.

  • Lesson 3 • Spark Performance Basics

    Introduces partitioning, caching, and broadcast joins for performance tuning. These techniques are applied in advanced optimization chapters later in the course.

  • Lesson 4 • Spark Architecture and Execution Model

    Explains drivers, executors, DAGs, and lazy evaluation. Understanding execution flow is critical for writing efficient Spark code in later chapters.

  • Lesson 5 • Spark SQL for Data Querying

    Teaches SQL-based querying on top of Spark using temporary views and catalog tables. Bridges SQL skills with the Spark execution engine.

Chapter 3See details

Delta Lake: Reliable Data Storage

  • Lesson 1 • DML Operations on Delta Tables

    Teaches INSERT, UPDATE, DELETE, and MERGE operations on Delta tables. Enables learners to implement upsert patterns essential for incremental data pipelines.

  • Lesson 2 • Optimizing Delta Tables

    Covers OPTIMIZE, ZORDER, and VACUUM commands for performance and storage management. These operations are critical for maintaining production-grade Delta tables.

  • Lesson 3 • Time Travel and Data Versioning

    Demonstrates querying historical table versions and restoring prior states. Time travel supports auditing and error recovery in production data pipelines.

  • Lesson 4 • Creating and Managing Delta Tables

    Covers DDL operations, table properties, and managed vs. external tables. These skills are required for building structured data layers in later pipeline chapters.

  • Lesson 5 • Delta Lake Core Concepts

    Explains the Delta transaction log, ACID guarantees, and how Delta differs from plain Parquet. Provides the conceptual foundation for all Delta table operations.

Chapter 4See details

Data Ingestion and ETL Pipelines

  • Lesson 1 • Auto Loader for Incremental Ingestion

    Introduces Auto Loader's file notification and directory listing modes for scalable ingestion. Auto Loader is the recommended Databricks pattern for continuous file-based ingestion.

  • Lesson 2 • Batch Ingestion Patterns

    Covers full and incremental batch loading strategies using Spark and Delta. Establishes the ingestion patterns that underpin all pipeline architectures in this chapter.

  • Lesson 3 • Building Transformation Logic

    Covers multi-step transformation patterns including cleansing, enrichment, and deduplication. Transformation logic is the core value-add layer in any ETL pipeline.

  • Lesson 4 • Delta Live Tables for Pipeline Orchestration

    Introduces Delta Live Tables (DLT) for declarative pipeline definition and quality enforcement. DLT simplifies pipeline management and integrates with the medallion architecture.

  • Lesson 5 • Structured Streaming Fundamentals

    Teaches micro-batch and continuous streaming with Spark Structured Streaming. Provides the streaming foundation needed for real-time pipeline sections ahead.

Chapter 5See details

Data Modeling and the Medallion Architecture

  • Lesson 1 • Medallion Architecture Principles

    Defines Bronze, Silver, and Gold layer responsibilities and data contracts. Provides the structural framework applied in every subsequent modeling and pipeline section.

  • Lesson 2 • Designing the Bronze Layer

    Covers raw data landing, schema-on-read, and preserving source fidelity. A well-designed Bronze layer ensures full auditability and reprocessing capability.

  • Lesson 3 • Constructing Gold Layer Aggregates

    Covers building pre-aggregated, denormalized tables optimized for BI and reporting. Gold tables reduce query complexity and improve dashboard performance.

  • Lesson 4 • Building the Silver Layer

    Teaches cleansing, type casting, deduplication, and conforming data to a unified schema. The Silver layer is the primary source for analytical and ML workloads.

  • Lesson 5 • Data Quality and Expectations

    Introduces constraint enforcement, quarantine patterns, and quality metrics tracking. Embedding quality checks at each layer prevents bad data from propagating downstream.

Chapter 6See details

Databricks SQL and Analytics

  • Lesson 1 • Query Optimization for Analytics

    Teaches query plan analysis, caching strategies, and result caching for SQL workloads. Optimized queries reduce warehouse costs and improve user experience.

  • Lesson 2 • Building Dashboards and Visualizations

    Covers creating charts, dashboards, and alerts within Databricks SQL. Dashboards enable business stakeholders to consume data without writing code.

  • Lesson 3 • Unity Catalog for Data Governance

    Introduces Unity Catalog's three-level namespace, access control, and data lineage. Governance is foundational for secure, compliant analytics across teams.

  • Lesson 4 • Advanced SQL Techniques in Databricks

    Teaches window functions, CTEs, and higher-order functions for complex analytical queries. These techniques enable sophisticated analysis directly within Databricks SQL.

  • Lesson 5 • SQL Warehouses and Query Execution

    Covers SQL warehouse types, sizing, and query routing. Proper warehouse configuration directly impacts query performance and cost for analytics workloads.

Chapter 7See details

Machine Learning with Databricks and MLflow

  • Lesson 1 • MLflow Model Registry

    Introduces model registration, stage transitions, and version management in the Model Registry. The registry provides governance and traceability for production models.

  • Lesson 2 • Model Serving and Batch Inference

    Covers real-time model serving endpoints and batch inference with Spark. Deploying models in both modes ensures flexibility for diverse business requirements.

  • Lesson 3 • Training Models at Scale with Spark

    Covers distributed training using pandas UDFs, Spark ML, and Hyperopt for tuning. Distributed training reduces time-to-insight on large datasets.

  • Lesson 4 • Experiment Tracking with MLflow

    Teaches logging parameters, metrics, and artifacts with MLflow Tracking. Systematic experiment tracking enables reproducibility and informed model selection.

  • Lesson 5 • Databricks ML Runtime and Feature Engineering

    Covers the ML Runtime environment, pre-installed libraries, and Feature Store basics. A well-prepared feature set is the foundation of reproducible model training.

Chapter 8See details

Workflow Orchestration and Production Deployment

  • Lesson 1 • Monitoring, Logging, and Cost Management

    Covers cluster event logs, job metrics, audit logs, and DBU cost tracking. Observability and cost control are essential for sustainable production operations.

  • Lesson 2 • Scheduling and Triggering Workflows

    Teaches cron-based scheduling, file-arrival triggers, and external API triggers. Flexible triggering ensures pipelines run at the right time with minimal manual intervention.

  • Lesson 3 • Infrastructure as Code with Terraform

    Introduces provisioning Databricks resources using the Terraform provider. Infrastructure as code ensures reproducible, auditable environment setup across workspaces.

  • Lesson 4 • Databricks Jobs and Task Orchestration

    Covers creating multi-task jobs, setting dependencies, and configuring compute for each task. Jobs are the primary mechanism for running production workloads on a schedule.

  • Lesson 5 • CI/CD for Databricks with Repos

    Covers Git integration via Databricks Repos, branch strategies, and automated deployment. CI/CD practices reduce deployment risk and accelerate delivery of data products.

Certification

Your valid completion certificate

This course is for you:

  • Data Engineer: ready to move beyond basic pipelines into a unified lakehouse platform.

  • Data Analyst: wants to query and visualize data at scale using Databricks SQL.

  • Data Scientist: needs a reliable platform for feature engineering and model deployment.

  • Analytics Engineer: building dbt-style transformation layers and exploring Spark-native alternatives.

  • Cloud Engineer: supporting data teams and needing deeper Databricks infrastructure knowledge.

  • Career Changer: transitioning from software development into the modern data engineering field.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQ

Who is Dedika?

Is the certificate valid in the United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course