Choose your language
Data Lake Course
Over 2 million learners across the globe

Data Lake Course

Master every layer of modern data lake engineering, from raw ingestion to governed, analytics-ready lakehouses. This course covers architecture, storage, processing, security, and DataOps in one comprehensive programme. Whether you're building your first lake or scaling an enterprise platform, you'll gain the hands-on expertise to deliver real business value.

Dedika for businesses

What you will learn:

You will learn how to design and operate production-grade data lakes, starting with core architecture concepts and advancing through storage formats, ingestion pipelines, metadata management, and governance. You will implement open table formats including Delta Lake, Apache Iceberg, and Apache Hudi to bring ACID transactions and time travel to lake storage. The course also covers distributed processing, query engine configuration, streaming architectures, and ML integration. You will apply DataOps practices such as CI/CD, infrastructure as code, and data contracts to automate pipeline delivery. Finally, you will learn cost management, observability, and how to communicate data lake value to executive stakeholders.

How you study practically Data Lake Course

How you practise Data Lake Course

For companies looking to train their teams

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Data Lake Architecture

  • Lesson 1 • What Is a Data Lake

    Defines data lakes and contrasts them with data warehouses and data marts. Provides the conceptual baseline for all subsequent architectural decisions.

  • Lesson 2 • Common Pitfalls and the Data Swamp Problem

    Identifies governance failures that turn data lakes into unusable data swamps. Awareness of these risks motivates best practices introduced later.

  • Lesson 3 • Data Lake Use Cases and Business Value

    Maps real-world scenarios—analytics, ML, archiving—to data lake capabilities. Learners connect technical features to measurable business outcomes.

  • Lesson 4 • Core Characteristics of Data Lakes

    Examines schema-on-read, multi-format ingestion, and scalability as defining traits. These properties explain why data lakes suit diverse analytical workloads.

Chapter 2See details

Data Lake Storage and File Formats

  • Lesson 1 • Storage Tiering and Cost Optimization

    Explains hot, warm, and cold storage tiers and automated tiering policies. Learners design cost-effective retention strategies aligned with access patterns.

  • Lesson 2 • Compression Strategies for Data Lakes

    Covers Snappy, Gzip, Zstd, and LZ4 trade-offs between compression ratio and speed. Proper compression reduces storage cost without degrading query performance.

  • Lesson 3 • Columnar and Row-Based File Formats

    Compares Parquet, ORC, Avro, CSV, and JSON for analytical workloads. Format choice directly impacts query speed and storage efficiency.

  • Lesson 4 • Object Storage as the Foundation

    Explains object storage mechanics, bucket organisation, and durability guarantees. Object storage is the dominant physical layer for modern data lakes.

  • Lesson 5 • Partitioning and File Sizing Best Practices

    Teaches partition key selection, partition pruning, and optimal file sizes. These techniques prevent small-file problems and accelerate analytical queries.

Chapter 3See details

Data Ingestion Patterns and Pipelines

  • Lesson 1 • Pipeline Reliability and Monitoring

    Teaches dead-letter queues, data quality checks at ingestion, and pipeline observability. Reliable pipelines prevent silent data loss and downstream corruption.

  • Lesson 2 • Streaming Ingestion Fundamentals

    Introduces event streaming concepts, message brokers, and exactly-once semantics. Streaming ingestion enables near-real-time analytics on continuously arriving data.

  • Lesson 3 • Ingestion Architecture Overview

    Maps ingestion patterns to source types and latency requirements. This overview frames the design decisions covered throughout the chapter.

  • Lesson 4 • Ingestion from Diverse Source Types

    Addresses databases, APIs, files, IoT devices, and SaaS platforms as ingestion sources. Each source type requires specific connectors and schema handling strategies.

  • Lesson 5 • Batch Ingestion Design

    Covers scheduled extraction, full vs. incremental loads, and change data capture. Batch ingestion remains the most common pattern for structured source systems.

Chapter 4See details

Data Lake Organisation and Metadata Management

  • Lesson 1 • Technical and Business Metadata

    Differentiates technical metadata (schema, lineage) from business metadata (ownership, glossary). Both types are required for effective data discovery and governance.

  • Lesson 2 • Zone-Based Lake Architecture

    Introduces raw, cleansed, curated, and sandbox zones with clear promotion criteria. Zone separation enforces data quality boundaries and access control.

  • Lesson 3 • Data Lineage Tracking

    Explains column-level and dataset-level lineage capture and visualisation. Lineage enables impact analysis and builds consumer trust in data assets.

  • Lesson 4 • Naming Conventions and Folder Structures

    Establishes standards for bucket names, prefixes, and dataset identifiers. Consistent naming enables automation, discovery, and governance at scale.

  • Lesson 5 • Data Catalogue Design and Implementation

    Covers catalogue architecture, crawlers, and search interfaces for dataset discovery. A well-designed catalogue reduces time-to-insight for data consumers.

Chapter 5See details

Data Governance and Security in Data Lakes

  • Lesson 1 • Access Control Models for Data Lakes

    Compares role-based, attribute-based, and row/column-level access control models. Selecting the right model balances security with operational flexibility.

  • Lesson 2 • Data Masking and Anonymization

    Covers static masking, dynamic masking, tokenization, and pseudonymization techniques. These methods protect sensitive data while preserving analytical utility.

  • Lesson 3 • Encryption at Rest and in Transit

    Explains server-side encryption, client-side encryption, and TLS for data in motion. Encryption is the baseline control for protecting lake data from unauthorized access.

  • Lesson 4 • Audit Logging and Compliance Reporting

    Designs audit log pipelines, retention policies, and compliance dashboards. Audit trails demonstrate control effectiveness to internal and external reviewers.

  • Lesson 5 • Data Classification and Sensitivity Tagging

    Defines classification tiers—public, internal, confidential, restricted—and tagging workflows. Classification drives downstream access and masking decisions.

Chapter 6See details

Data Processing and Transformation at Scale

  • Lesson 1 • Distributed Processing Fundamentals

    Introduces distributed execution models, DAGs, and shuffle operations. Understanding these mechanics is prerequisite to optimizing any large-scale transformation job.

  • Lesson 2 • Data Quality Transformation Rules

    Implements validation, standardization, deduplication, and enrichment as transformation steps. Embedding quality rules in pipelines prevents bad data from reaching consumers.

  • Lesson 3 • Pipeline Performance Optimization

    Applies caching, partition pruning, predicate pushdown, and job tuning to reduce runtime. Optimised pipelines lower compute costs and meet SLA requirements.

  • Lesson 4 • Streaming Transformation Patterns

    Teaches stateful stream processing, watermarks, and late-arriving data handling. Streaming transformations enable continuous analytics on real-time data flows.

  • Lesson 5 • Batch Transformation Patterns

    Covers filter, join, aggregate, and window operations on large datasets. These patterns form the building blocks of all analytical transformation pipelines.

Chapter 7See details

Open Table Formats and Lakehouse Architecture

  • Lesson 1 • Lakehouse Architecture Design

    Integrates open table formats with query engines, BI tools, and ML platforms into a unified lakehouse. This architecture eliminates the need for a separate data warehouse in many scenarios.

  • Lesson 2 • Apache Iceberg and Apache Hudi

    Compares Iceberg's hidden partitioning and Hudi's record-level upserts to Delta Lake. Learners select the right format based on workload and ecosystem requirements.

  • Lesson 3 • Limitations of Traditional Data Lakes

    Identifies consistency, update, and time-travel gaps in plain object-store lakes. These limitations motivate the adoption of open table formats.

  • Lesson 4 • Open Table Format Concepts

    Explains metadata layers, transaction logs, and snapshot isolation common to all open formats. A shared conceptual model simplifies comparison of specific implementations.

  • Lesson 5 • Delta Lake Deep Dive

    Covers Delta log mechanics, MERGE operations, Z-ordering, and vacuum commands. Delta Lake is widely adopted and serves as the primary implementation reference.

Chapter 8See details

Data Lake Operations and Strategic Governance

  • Lesson 1 • Data Retention and Lifecycle Governance

    Designs retention schedules, deletion workflows, and legal hold processes. Lifecycle governance reduces risk and storage cost while satisfying regulatory obligations.

  • Lesson 2 • Data Mesh and Federated Ownership

    Applies data mesh principles—domain ownership, data products, and federated governance—to large lakes. Federated models scale governance beyond central team capacity.

  • Lesson 3 • Roadmap Planning and Maturity Assessment

    Uses maturity models to assess current lake capabilities and prioritize improvement initiatives. A structured roadmap aligns technical investments with business strategy.

  • Lesson 4 • Cost Management and FinOps for Data Lakes

    Applies tagging, showback, chargeback, and rightsizing to control lake spending. Cost visibility drives accountability and prevents uncontrolled storage and compute growth.

  • Lesson 5 • Observability and Monitoring

    Builds monitoring stacks covering pipeline health, data freshness, and storage metrics. Observability enables proactive issue resolution before consumers are impacted.

Certification

Your valid completion certificate

This course is for you:

  • Data engineers who want to move beyond basic pipeline work.

  • Cloud architects designing scalable storage solutions for analytics teams.

  • BI developers whose reporting needs outgrow traditional warehouse limitations.

  • Software engineers transitioning into data infrastructure and platform roles.

  • IT managers overseeing data strategy and needing technical depth fast.

  • Analytics engineers building reproducible ML-ready datasets from raw sources.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.
André Felipe
André FelipePrompt Engineering Student

Top training programmes

FAQ

Who is Dedika?

Is the certificate valid in Kenya?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course