
Data Lake Course
Master every layer of modern data lake engineering, from raw ingestion to governed, analytics-ready lakehouses. This course covers architecture, storage, processing, security, and DataOps in one comprehensive programme. Whether you're building your first lake or scaling an enterprise platform, you'll gain the hands-on expertise to deliver real business value.
What you will learn:
You will learn how to design and operate production-grade data lakes, starting with core architecture concepts and advancing through storage formats, ingestion pipelines, metadata management, and governance. You will implement open table formats including Delta Lake, Apache Iceberg, and Apache Hudi to bring ACID transactions and time travel to lake storage. The course also covers distributed processing, query engine configuration, streaming architectures, and ML integration. You will apply DataOps practices such as CI/CD, infrastructure as code, and data contracts to automate pipeline delivery. Finally, you will learn cost management, observability, and how to communicate data lake value to executive stakeholders.
How you study in practice Data Lake Course
How you practise Data Lake Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Data Lake Architecture
Foundations of Data Lake Architecture
Lesson 1 • What Is a Data Lake
Defines data lakes and contrasts them with data warehouses and data marts. Provides the conceptual baseline for all subsequent architectural decisions.
Lesson 2 • Common Pitfalls and the Data Swamp Problem
Identifies governance failures that turn data lakes into unusable data swamps. Awareness of these risks motivates best practices introduced later.
Lesson 3 • Data Lake Use Cases and Business Value
Maps real-world scenarios—analytics, ML, archiving—to data lake capabilities. Learners connect technical features to measurable business outcomes.
Lesson 4 • Core Characteristics of Data Lakes
Examines schema-on-read, multi-format ingestion, and scalability as defining traits. These properties explain why data lakes suit diverse analytical workloads.
Chapter 2HideHide detailsSee detailsData Lake Storage and File Formats
Data Lake Storage and File Formats
Lesson 1 • Storage Tiering and Cost Optimization
Explains hot, warm, and cold storage tiers and automated tiering policies. Learners design cost-effective retention strategies aligned with access patterns.
Lesson 2 • Compression Strategies for Data Lakes
Covers Snappy, Gzip, Zstd, and LZ4 trade-offs between compression ratio and speed. Proper compression reduces storage cost without degrading query performance.
Lesson 3 • Columnar and Row-Based File Formats
Compares Parquet, ORC, Avro, CSV, and JSON for analytical workloads. Format choice directly impacts query speed and storage efficiency.
Lesson 4 • Object Storage as the Foundation
Explains object storage mechanics, bucket organisation, and durability guarantees. Object storage is the dominant physical layer for modern data lakes.
Lesson 5 • Partitioning and File Sizing Best Practices
Teaches partition key selection, partition pruning, and optimal file sizes. These techniques prevent small-file problems and accelerate analytical queries.
Chapter 3HideHide detailsSee detailsData Ingestion Patterns and Pipelines
Data Ingestion Patterns and Pipelines
Lesson 1 • Pipeline Reliability and Monitoring
Teaches dead-letter queues, data quality checks at ingestion, and pipeline observability. Reliable pipelines prevent silent data loss and downstream corruption.
Lesson 2 • Streaming Ingestion Fundamentals
Introduces event streaming concepts, message brokers, and exactly-once semantics. Streaming ingestion enables near-real-time analytics on continuously arriving data.
Lesson 3 • Ingestion Architecture Overview
Maps ingestion patterns to source types and latency requirements. This overview frames the design decisions covered throughout the chapter.
Lesson 4 • Ingestion from Diverse Source Types
Addresses databases, APIs, files, IoT devices, and SaaS platforms as ingestion sources. Each source type requires specific connectors and schema handling strategies.
Lesson 5 • Batch Ingestion Design
Covers scheduled extraction, full vs. incremental loads, and change data capture. Batch ingestion remains the most common pattern for structured source systems.
Chapter 4HideHide detailsSee detailsData Lake Organisation and Metadata Management
Data Lake Organisation and Metadata Management
Lesson 1 • Technical and Business Metadata
Differentiates technical metadata (schema, lineage) from business metadata (ownership, glossary). Both types are required for effective data discovery and governance.
Lesson 2 • Zone-Based Lake Architecture
Introduces raw, cleansed, curated, and sandbox zones with clear promotion criteria. Zone separation enforces data quality boundaries and access control.
Lesson 3 • Data Lineage Tracking
Explains column-level and dataset-level lineage capture and visualisation. Lineage enables impact analysis and builds consumer trust in data assets.
Lesson 4 • Naming Conventions and Folder Structures
Establishes standards for bucket names, prefixes, and dataset identifiers. Consistent naming enables automation, discovery, and governance at scale.
Lesson 5 • Data Catalogue Design and Implementation
Covers catalogue architecture, crawlers, and search interfaces for dataset discovery. A well-designed catalogue reduces time-to-insight for data consumers.
Chapter 5HideHide detailsSee detailsData Governance and Security in Data Lakes
Data Governance and Security in Data Lakes
Lesson 1 • Access Control Models for Data Lakes
Compares role-based, attribute-based, and row/column-level access control models. Selecting the right model balances security with operational flexibility.
Lesson 2 • Data Masking and Anonymization
Covers static masking, dynamic masking, tokenization, and pseudonymization techniques. These methods protect sensitive data while preserving analytical utility.
Lesson 3 • Encryption at Rest and in Transit
Explains server-side encryption, client-side encryption, and TLS for data in motion. Encryption is the baseline control for protecting lake data from unauthorized access.
Lesson 4 • Audit Logging and Compliance Reporting
Designs audit log pipelines, retention policies, and compliance dashboards. Audit trails demonstrate control effectiveness to internal and external reviewers.
Lesson 5 • Data Classification and Sensitivity Tagging
Defines classification tiers—public, internal, confidential, restricted—and tagging workflows. Classification drives downstream access and masking decisions.
Chapter 6HideHide detailsSee detailsData Processing and Transformation at Scale
Data Processing and Transformation at Scale
Lesson 1 • Distributed Processing Fundamentals
Introduces distributed execution models, DAGs, and shuffle operations. Understanding these mechanics is prerequisite to optimizing any large-scale transformation job.
Lesson 2 • Data Quality Transformation Rules
Implements validation, standardization, deduplication, and enrichment as transformation steps. Embedding quality rules in pipelines prevents bad data from reaching consumers.
Lesson 3 • Pipeline Performance Optimization
Applies caching, partition pruning, predicate pushdown, and job tuning to reduce runtime. Optimised pipelines lower compute costs and meet SLA requirements.
Lesson 4 • Streaming Transformation Patterns
Teaches stateful stream processing, watermarks, and late-arriving data handling. Streaming transformations enable continuous analytics on real-time data flows.
Lesson 5 • Batch Transformation Patterns
Covers filter, join, aggregate, and window operations on large datasets. These patterns form the building blocks of all analytical transformation pipelines.
Chapter 7HideHide detailsSee detailsOpen Table Formats and Lakehouse Architecture
Open Table Formats and Lakehouse Architecture
Lesson 1 • Lakehouse Architecture Design
Integrates open table formats with query engines, BI tools, and ML platforms into a unified lakehouse. This architecture eliminates the need for a separate data warehouse in many scenarios.
Lesson 2 • Apache Iceberg and Apache Hudi
Compares Iceberg's hidden partitioning and Hudi's record-level upserts to Delta Lake. Learners select the right format based on workload and ecosystem requirements.
Lesson 3 • Limitations of Traditional Data Lakes
Identifies consistency, update, and time-travel gaps in plain object-store lakes. These limitations motivate the adoption of open table formats.
Lesson 4 • Open Table Format Concepts
Explains metadata layers, transaction logs, and snapshot isolation common to all open formats. A shared conceptual model simplifies comparison of specific implementations.
Lesson 5 • Delta Lake Deep Dive
Covers Delta log mechanics, MERGE operations, Z-ordering, and vacuum commands. Delta Lake is widely adopted and serves as the primary implementation reference.
Chapter 8HideHide detailsSee detailsData Lake Operations and Strategic Governance
Data Lake Operations and Strategic Governance
Lesson 1 • Data Retention and Lifecycle Governance
Designs retention schedules, deletion workflows, and legal hold processes. Lifecycle governance reduces risk and storage cost while satisfying regulatory obligations.
Lesson 2 • Data Mesh and Federated Ownership
Applies data mesh principles—domain ownership, data products, and federated governance—to large lakes. Federated models scale governance beyond central team capacity.
Lesson 3 • Roadmap Planning and Maturity Assessment
Uses maturity models to assess current lake capabilities and prioritize improvement initiatives. A structured roadmap aligns technical investments with business strategy.
Lesson 4 • Cost Management and FinOps for Data Lakes
Applies tagging, showback, chargeback, and rightsizing to control lake spending. Cost visibility drives accountability and prevents uncontrolled storage and compute growth.
Lesson 5 • Observability and Monitoring
Builds monitoring stacks covering pipeline health, data freshness, and storage metrics. Observability enables proactive issue resolution before consumers are impacted.
Your valid completion certificate
This course is for you:
Data engineers who want to move beyond basic pipeline work.
Cloud architects designing scalable storage solutions for analytics teams.
BI developers whose reporting needs outgrow traditional warehouse limitations.
Software engineers transitioning into data infrastructure and platform roles.
IT managers overseeing data strategy and needing technical depth fast.
Analytics engineers building reproducible ML-ready datasets from raw sources.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















