
Data Engineering Course
Master the full data engineering stack — from SQL and Python to Kafka, Spark, and dbt. This course takes you through every layer of modern data infrastructure, giving you the hands-on skills employers are actively hiring for. Build real pipelines, work with industry-standard tools, and come out ready to deliver on day one.
What you will learn:
You will learn how to design and implement end-to-end data pipelines that move, transform, and serve data at scale. The course covers relational modelling, advanced SQL, Python scripting, and cloud data warehouse engineering. You will work with tools like Apache Airflow, dbt, Apache Kafka, and Apache Spark. Topics include data ingestion patterns, streaming architectures, DataOps practices, and data governance frameworks. By the end, you will have the technical depth and practical experience to operate as a professional data engineer in any modern data-driven organisation.
How you study practically Data Engineering Course
How you practise Data Engineering Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 37 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Data Engineering
Foundations of Data Engineering
Lesson 1 • Introduction to Data Architecture Patterns
Surveys foundational architectural styles including data warehouses, data lakes, and lakehouses. Connects architectural choices to business requirements and scalability goals.
Lesson 2 • Core Data Concepts and Terminology
Introduces structured, semi-structured, and unstructured data types and key storage abstractions. Ensures shared vocabulary before diving into architecture and tooling.
Lesson 3 • Data Engineering Lifecycle Overview
Maps the end-to-end journey of data from source to consumption. Provides a mental model that frames every technical topic covered later in the course.
Lesson 4 • The Data Engineering Role
Defines the data engineer's responsibilities within data teams and business contexts. Establishes the professional baseline needed for all subsequent technical chapters.
Chapter 2HideHide detailsSee detailsSQL and Relational Data Modeling
SQL and Relational Data Modeling
Lesson 1 • Advanced SQL for Data Engineering
Introduces window functions, CTEs, and set operations for complex analytical transformations. Prepares learners to write production-grade SQL used in pipelines and reporting layers.
Lesson 2 • Dimensional Modeling for Analytics
Applies star and snowflake schema patterns to design analytics-ready data models. Directly enables the data warehouse and transformation work covered in later chapters.
Lesson 3 • Core SQL Querying Techniques
Teaches SELECT, filtering, aggregation, and multi-table joins for data retrieval. These skills underpin every transformation and analytical workload in the course.
Lesson 4 • Relational Database Fundamentals
Covers tables, keys, constraints, and relationships as the building blocks of relational systems. Grounds learners in the theory required for effective schema design.
Chapter 3HideHide detailsSee detailsPython for Data Engineering
Python for Data Engineering
Lesson 1 • Python Essentials for Data Work
Reviews Python syntax, data structures, and control flow relevant to data manipulation tasks. Establishes the programming foundation required for pipeline scripting throughout the course.
Lesson 2 • File and Data Format Handling
Teaches reading and writing CSV, JSON, Parquet, and Avro files using Python libraries. Directly supports ingestion and serialisation tasks in pipeline development chapters.
Lesson 3 • Automation and Scripting Best Practices
Introduces logging, configuration management, and packaging for production-ready Python scripts. Prepares learners to write maintainable code that integrates into orchestrated pipelines.
Lesson 4 • Data Manipulation with Pandas
Covers DataFrame operations, merging, reshaping, and cleaning using the Pandas library. Provides the transformation toolkit used in batch processing and exploratory data work.
Lesson 5 • Connecting Python to Databases and APIs
Demonstrates database connectivity via SQL drivers and REST API consumption using Python. Bridges scripting skills to real-world data source integration covered in ingestion chapters.
Chapter 4HideHide detailsSee detailsData Ingestion and Pipeline Design
Data Ingestion and Pipeline Design
Lesson 1 • Extracting Data from Relational Sources
Covers CDC, timestamp-based extraction, and query-based pulls from relational databases. Builds on SQL skills to implement reliable, low-impact extraction strategies.
Lesson 2 • Pipeline Design Principles
Introduces modularity, idempotency, error handling, and observability as pipeline design pillars. Establishes engineering standards applied throughout all subsequent pipeline chapters.
Lesson 3 • Ingestion Patterns and Source Types
Surveys common data sources including databases, APIs, files, and event streams. Frames the ingestion design decisions that drive pipeline architecture choices.
Lesson 4 • Ingesting from APIs and File Sources
Teaches structured approaches to consuming REST APIs, FTP sources, and cloud storage files. Extends Python API skills into production ingestion pipeline patterns.
Lesson 5 • Data Validation and Quality Checks
Implements schema validation, null checks, and statistical profiling within ingestion pipelines. Ensures data quality is enforced at the point of entry before downstream processing.
Chapter 5HideHide detailsSee detailsData Storage and Warehouse Engineering
Data Storage and Warehouse Engineering
Lesson 1 • Table Formats and Open Storage Standards
Introduces open table formats that enable ACID transactions and schema evolution on object storage. Prepares learners for lakehouse implementations covered in advanced chapters.
Lesson 2 • Cloud Data Warehouse Architecture
Examines columnar storage, MPP query engines, and separation of compute and storage. Connects warehouse internals to query performance and cost optimisation decisions.
Lesson 3 • Data Modeling in the Warehouse
Applies dimensional and vault modeling techniques within cloud warehouse environments. Builds on relational modeling skills to produce analytics-ready warehouse layers.
Lesson 4 • Cloud Object Storage Fundamentals
Explains object storage concepts, bucket organization, partitioning, and access control. Provides the storage foundation underlying data lakes and lakehouse architectures.
Lesson 5 • Warehouse Performance and Cost Optimization
Covers query profiling, materialized views, caching, and warehouse sizing strategies. Equips learners to balance performance and cost in production warehouse environments.
Chapter 6HideHide detailsSee detailsData Transformation and dbt
Data Transformation and dbt
Lesson 1 • Advanced dbt Patterns
Covers macros, Jinja templating, packages, and incremental strategies for complex use cases. Extends core dbt skills to handle large-scale and dynamic transformation requirements.
Lesson 2 • Testing and Documentation in dbt
Implements schema tests, custom tests, and auto-generated documentation within dbt projects. Enforces data quality and maintainability standards in transformation pipelines.
Lesson 3 • Transformation Layer Architecture
Defines staging, intermediate, and mart transformation layers and their responsibilities. Establishes the layered design pattern that structures all dbt project work in this chapter.
Lesson 4 • Building dbt Models
Teaches writing SQL models, configuring materializations, and managing model dependencies in dbt. Translates SQL skills into a version-controlled, modular transformation workflow.
Chapter 7HideHide detailsSee detailsPipeline Orchestration and Workflow Management
Pipeline Orchestration and Workflow Management
Lesson 1 • Pipeline Monitoring and Observability
Establishes logging, metrics, and lineage tracking practices for orchestrated pipeline visibility. Connects observability tooling to the data quality and governance themes of the course.
Lesson 2 • Modern Orchestration Alternatives
Surveys Prefect, Dagster, and event-driven orchestration as alternatives to Airflow. Equips learners to evaluate orchestration tools against project requirements and team constraints.
Lesson 3 • Error Handling and Retry Strategies
Implements retries, SLAs, alerting, and failure callbacks within orchestrated workflows. Ensures pipelines meet reliability standards required in production data environments.
Lesson 4 • Orchestration Concepts and Patterns
Introduces DAG-based scheduling, dependency resolution, and orchestration vs. execution separation. Frames the orchestration layer that ties ingestion, storage, and transformation together.
Lesson 5 • Building Workflows with Apache Airflow
Covers Airflow architecture, operators, hooks, and DAG authoring in Python. Provides hands-on experience with the industry-standard orchestration tool for data pipelines.
Chapter 8HideHide detailsSee detailsStreaming Data Engineering
Streaming Data Engineering
Lesson 1 • Streaming Data Storage and Serving
Addresses writing streaming outputs to data lakes, warehouses, and real-time serving stores. Connects streaming pipelines to the storage and consumption layers built in earlier chapters.
Lesson 2 • Streaming Pipeline Reliability and Scaling
Covers backpressure, schema registry, dead-letter topics, and horizontal scaling for streaming systems. Ensures learners can operate streaming pipelines reliably at production scale.
Lesson 3 • Stream Processing with Apache Flink
Teaches stateful stream processing, windowing, and event-time handling using Apache Flink. Builds on Kafka knowledge to implement real-time transformation and aggregation logic.
Lesson 4 • Streaming Fundamentals and Use Cases
Contrasts streaming with batch processing and identifies use cases where real-time data is required. Establishes the conceptual foundation for all streaming architecture and tooling decisions.
Lesson 5 • Apache Kafka Architecture and Operations
Covers Kafka brokers, topics, partitions, consumer groups, and producer/consumer APIs. Provides the messaging backbone knowledge required for all streaming pipeline implementations.
Your valid completion certificate
This course is for you:
Business analyst: ready to move from consuming data to engineering it.
BI developer: wants to own the pipeline layer behind the dashboards.
Software developer: looking to specialise in data infrastructure and tooling.
Recent graduate: entering the job market with a focus on data roles.
Data analyst: aiming to level up into a higher-impact engineering position.
Career changer: coming from IT or operations and pivoting into data engineering.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.

Top training programmes
FAQ
Who is Dedika?
Is the certificate valid in Kenya?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















