Choose your language
Data Engineering with Databricks Course
More than 2 million students worldwide

Data Engineering with Databricks Course

Master the full Databricks data engineering stack — from Apache Spark fundamentals and Delta Lake architecture to production-grade pipeline orchestration and enterprise governance. This course equips you with the hands-on skills to design, optimize, and operate reliable data platforms that scale. Whether you're breaking into data engineering or leveling up your lakehouse expertise, this is the definitive Databricks training.

Dedika for Business

What you will learn:

  • Configure and manage Databricks clusters, workspaces, and compute resources efficiently.

  • Build scalable ELT pipelines using Auto Loader, Structured Streaming, and medallion architecture.

  • Implement Delta Lake tables with ACID transactions, time travel, and schema evolution.

  • Create declarative, self-healing pipelines with Delta Live Tables and quality expectations.

  • Apply advanced Spark tuning techniques including AQE, partitioning strategies, and Photon engine.

  • Enforce enterprise data governance, access control, and lineage tracking using Unity Catalog.

How you study in practice Data Engineering with Databricks Course

How you practise Data Engineering with Databricks Course

For companies looking to train their team

With Dedika for Business, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course Content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Databricks Platform Foundations

  • Lesson 1 • Cluster Configuration and Management

    Teaches cluster creation, sizing, autoscaling, and termination policies. Students gain control over compute costs and performance from the start.

  • Lesson 2 • Databricks Architecture Overview

    Covers the control plane, data plane, and cluster architecture underpinning Databricks. Grounds all subsequent platform usage in a clear mental model.

  • Lesson 3 • Workspace Navigation and Setup

    Introduces the Databricks UI, folder structure, and user settings. Enables students to organize projects and manage personal environments effectively.

  • Lesson 4 • Databricks CLI and REST API Basics

    Introduces programmatic access to the platform via CLI and REST API. Prepares students for automation tasks covered in later chapters.

  • Lesson 5 • Notebooks and Collaborative Development

    Explores notebook creation, magic commands, and real-time collaboration features. Establishes the primary development interface used throughout the course.

Chapter 2See details

Apache Spark Core Concepts

  • Lesson 1 • Transformations and Actions

    Distinguishes narrow vs. wide transformations and explains action triggers. Students learn to reason about shuffle costs and execution boundaries.

  • Lesson 2 • Spark Distributed Execution Model

    Explains RDDs, DAGs, stages, and tasks within Spark's execution engine. Provides the conceptual foundation for writing efficient distributed code.

  • Lesson 3 • DataFrame API Fundamentals

    Covers DataFrame creation, schema inference, and core transformations. Students apply these operations as the primary data manipulation interface.

  • Lesson 4 • Spark Performance Fundamentals

    Introduces the Spark UI, execution plans, and basic tuning levers. Equips students to diagnose and address common performance bottlenecks.

  • Lesson 5 • Spark SQL and Catalog

    Teaches SQL queries on DataFrames and the Spark catalog for metadata management. Bridges SQL knowledge with programmatic Spark usage.

Chapter 3See details

Delta Lake Architecture and Operations

  • Lesson 1 • Delta Lake Storage Format

    Explains Parquet files, the transaction log, and checkpoint files that form Delta Lake. Students understand why Delta provides reliability over raw file formats.

  • Lesson 2 • Creating and Managing Delta Tables

    Covers DDL for creating, altering, and dropping Delta tables. Establishes the table management skills required for all pipeline work ahead.

  • Lesson 3 • ACID Transactions and Concurrency

    Teaches optimistic concurrency control, isolation levels, and conflict resolution in Delta. Students can design pipelines that safely handle concurrent writes.

  • Lesson 4 • Delta Lake Optimization Techniques

    Covers OPTIMIZE, Z-ordering, and file compaction to improve query performance. Students apply these techniques to production-grade Delta tables.

  • Lesson 5 • Time Travel and Data Versioning

    Demonstrates querying historical table versions and restoring prior states. Enables audit, rollback, and reproducibility patterns in data pipelines.

Chapter 4See details

Data Ingestion and ELT Patterns

  • Lesson 1 • ELT Design Patterns

    Presents medallion architecture and incremental load strategies for structured ELT. Students design multi-hop pipelines that separate raw, cleansed, and curated layers.

  • Lesson 2 • Batch Ingestion from Common Sources

    Teaches reading from cloud object storage, JDBC databases, and file formats. Provides the ingestion foundation for all pipeline patterns in this chapter.

  • Lesson 3 • Data Quality Enforcement at Ingestion

    Teaches constraint-based validation, quarantine patterns, and expectation frameworks. Ensures data quality is enforced before data reaches downstream consumers.

  • Lesson 4 • Structured Streaming Fundamentals

    Covers streaming DataFrames, triggers, and output modes for continuous data processing. Connects streaming concepts to the Delta Lake sink used in production.

  • Lesson 5 • Auto Loader for Incremental Ingestion

    Introduces Auto Loader's file notification and directory listing modes for scalable incremental loads. Students configure Auto Loader checkpointing and schema evolution.

Chapter 5See details

Delta Live Tables and Pipeline Orchestration

  • Lesson 1 • Pipeline Configuration and Deployment

    Teaches pipeline settings, cluster policies, and deployment via UI and API. Students configure production-ready pipelines with appropriate compute and storage settings.

  • Lesson 2 • Monitoring and Troubleshooting DLT

    Explores the event log, pipeline UI metrics, and common failure patterns. Students diagnose and resolve pipeline issues using built-in observability tools.

  • Lesson 3 • Data Quality with Expectations

    Covers EXPECT, EXPECT OR DROP, and EXPECT OR FAIL constraints in DLT. Students embed quality rules directly into pipeline definitions for automated enforcement.

  • Lesson 4 • Orchestrating Pipelines with Databricks Workflows

    Integrates DLT pipelines into multi-task Databricks Workflows with dependencies and scheduling. Students build end-to-end orchestrated data products.

  • Lesson 5 • Delta Live Tables Core Concepts

    Introduces DLT syntax, live tables, streaming live tables, and the pipeline graph. Establishes the declarative paradigm that distinguishes DLT from imperative Spark code.

Chapter 6See details

Unity Catalog and Data Governance

  • Lesson 1 • Identity and Access Management

    Covers users, groups, service principals, and privilege grants in Unity Catalog. Students implement least-privilege access models for data assets.

  • Lesson 2 • Unity Catalog Architecture

    Explains the metastore, catalog, schema, and table hierarchy in Unity Catalog. Provides the structural understanding needed to design governed namespaces.

  • Lesson 3 • Sharing Data with Delta Sharing

    Introduces the open Delta Sharing protocol for cross-platform and cross-organization data sharing. Students publish and consume shared datasets securely.

  • Lesson 4 • Auditing and Compliance

    Teaches audit log access, query history, and compliance reporting patterns. Students configure monitoring to satisfy enterprise audit requirements.

  • Lesson 5 • Data Lineage and Discovery

    Demonstrates automatic lineage capture and the data explorer for asset discovery. Students use lineage to trace data flow and support impact analysis.

Chapter 7See details

Advanced Spark Optimization and Tuning

  • Lesson 1 • Cost-Based Optimization and Statistics

    Covers ANALYZE TABLE, column statistics, and the cost-based optimizer's join reordering. Students collect and leverage statistics to improve plan quality.

  • Lesson 2 • Memory Management and Spill

    Covers Spark memory regions, executor sizing, and spill detection and mitigation. Students configure memory settings to eliminate costly disk spill.

  • Lesson 3 • Photon Engine and Vectorized Execution

    Introduces Databricks Photon engine and its vectorized execution benefits. Students identify workloads that benefit most from Photon and enable it appropriately.

  • Lesson 4 • Partitioning and Shuffling Strategies

    Teaches optimal partition counts, repartition vs. coalesce, and shuffle tuning. Students eliminate data skew and reduce shuffle overhead in complex pipelines.

  • Lesson 5 • Adaptive Query Execution

    Explains AQE's runtime plan adjustments for joins, skew, and coalescing. Students enable and configure AQE to reduce manual tuning overhead.

Chapter 8See details

Production Data Engineering Practices

  • Lesson 1 • Cost Management and FinOps

    Covers DBU consumption analysis, cluster right-sizing, and spot instance strategies. Students reduce compute costs without sacrificing pipeline reliability.

  • Lesson 2 • CI/CD for Databricks Projects

    Teaches Databricks Asset Bundles, deployment automation, and environment promotion. Students implement a full CI/CD pipeline from development to production.

  • Lesson 3 • Security Hardening and Best Practices

    Teaches network isolation, secret management, and secure credential handling. Students apply defense-in-depth principles to protect data assets in production.

  • Lesson 4 • Testing Data Pipelines

    Covers unit testing with pytest, integration testing strategies, and test data management. Students build a test suite that validates pipeline logic before deployment.

  • Lesson 5 • Observability and Pipeline Monitoring

    Introduces logging frameworks, metrics emission, and alerting for data pipelines. Students instrument pipelines to detect failures and SLA breaches proactively.

Certification

Your valid completion certificate

This course is for you:

  • Data analysts: ready to shift from querying data to engineering reliable pipelines.

  • Backend developers: looking to pivot their coding skills into the data platform space.

  • Junior data engineers: seeking a structured path to production-level Databricks expertise.

  • BI developers: wanting to modernize their ETL work using lakehouse-native tooling.

  • DevOps engineers: expanding their scope to include data infrastructure and pipeline automation.

  • Graduate students: building job-ready data engineering skills before entering the workforce.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top training programs

FAQ

Who is Dedika?

Is the certificate valid in Canada?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course