Choose your language
Cloud Computing and Big Data Course
More than 2 million students worldwide

Cloud Computing and Big Data Course

Master cloud computing and big data engineering with a curriculum that takes you from core cloud concepts to production-grade data pipelines and machine learning deployment. You'll gain hands-on skills in Apache Spark, cloud-native orchestration, and scalable analytics. This course is built for professionals who want to design, build, and operate real data systems in the cloud.

Dedika for businesses

What you will learn:

You will learn how cloud service models, deployment architectures, and pricing structures work so you can make informed infrastructure decisions. You will build data ingestion pipelines that handle both batch and real-time streaming workloads at scale. You will develop practical skills in Apache Spark, Apache Airflow, and cloud-native managed services. You will apply big data analytics techniques and create interactive dashboards that communicate insights to business stakeholders. You will also explore MLOps practices, including feature stores, model serving, and automated retraining pipelines. Finally, you will cover cloud security, cost optimization with FinOps, and data governance frameworks that are required in professional environments.

How you study in practice Cloud Computing and Big Data Course

How you practice Cloud Computing and Big Data Course

For companies that want to train their team

With Dedika for Business, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Cloud Computing

  • Lesson 1 • Cloud Computing Concepts and History

    Traces cloud evolution from mainframes to modern hyperscalers and defines essential terminology. Provides the conceptual baseline for all subsequent cloud topics.

  • Lesson 2 • Cloud Service Models Explained

    Differentiates IaaS, PaaS, and SaaS through functional comparisons and real-world examples. Enables students to map workloads to the correct service layer.

  • Lesson 3 • Cloud Deployment Models

    Examines public, private, hybrid, and multi-cloud architectures with trade-off analysis. Connects deployment choice to organizational compliance and cost goals.

  • Lesson 4 • Cloud Economics and Business Value

    Covers CapEx-to-OpEx shift, pay-as-you-go pricing, and total cost of ownership analysis. Grounds technical decisions in measurable financial outcomes.

Chapter 2See details

Core Cloud Infrastructure and Services

  • Lesson 1 • Cloud Storage Services

    Covers object, block, and file storage types with durability and access-pattern trade-offs. Prepares students to select storage tiers for diverse application requirements.

  • Lesson 2 • Identity and Access Management

    Teaches role-based access control, policies, and least-privilege principles for cloud environments. Establishes security hygiene required before deploying any production workload.

  • Lesson 3 • Compute Services and Virtual Machines

    Introduces virtual machines, instance types, and auto-scaling groups as core compute primitives. Directly applies cloud service model knowledge to infrastructure provisioning.

  • Lesson 4 • Managed Database Services

    Surveys relational and NoSQL managed database offerings, covering provisioning and backup strategies. Connects storage knowledge to stateful application architecture.

  • Lesson 5 • Cloud Networking Fundamentals

    Explains virtual networks, subnets, routing, and load balancing as cloud networking building blocks. Enables secure and performant connectivity between cloud resources.

Chapter 3See details

Big Data Concepts and Ecosystem

  • Lesson 1 • Batch and Stream Processing Models

    Contrasts batch processing with real-time stream processing using the Lambda and Kappa architectures. Prepares students to design pipelines matching latency and throughput requirements.

  • Lesson 2 • Introduction to Hadoop Ecosystem

    Explains HDFS, MapReduce, YARN, and key ecosystem tools as the foundation of distributed batch processing. Provides historical and technical context for modern big data frameworks.

  • Lesson 3 • Big Data Storage Architectures

    Covers distributed file systems, data lakes, and data warehouses as complementary storage paradigms. Bridges cloud storage knowledge to big data ingestion and retention needs.

  • Lesson 4 • Big Data Pipeline Design Principles

    Introduces ingestion, transformation, storage, and serving layers as a unified pipeline framework. Gives students a design template applied throughout the remaining big data chapters.

  • Lesson 5 • Defining Big Data and Its Characteristics

    Defines the five Vs—volume, velocity, variety, veracity, and value—with industry examples. Establishes the vocabulary and scope for all subsequent big data chapters.

Chapter 4See details

Data Ingestion and Storage at Scale

  • Lesson 1 • Data Warehouse Loading Strategies

    Covers dimensional modeling, slowly changing dimensions, and bulk-load patterns for analytical stores. Connects ingestion pipelines to query-optimized warehouse structures.

  • Lesson 2 • Real-Time Data Streaming Ingestion

    Teaches message queues, event streaming platforms, and producer-consumer patterns for real-time data. Extends batch ingestion knowledge to low-latency, high-throughput scenarios.

  • Lesson 3 • Batch Ingestion Techniques

    Covers ETL workflows, file-based ingestion, and scheduled data transfers for large datasets. Applies pipeline design principles to structured and semi-structured data sources.

  • Lesson 4 • Data Quality and Validation

    Introduces data profiling, schema validation, and anomaly detection as ingestion quality gates. Ensures downstream analytics and ML models receive trustworthy data.

  • Lesson 5 • Data Lake Design and Organization

    Explains zone-based data lake organization, partitioning strategies, and metadata management. Ensures ingested data is discoverable, governed, and ready for downstream processing.

Chapter 5See details

Distributed Processing with Apache Spark

  • Lesson 1 • Spark Performance Tuning

    Addresses partitioning, caching, broadcast joins, and serialization as key Spark optimization levers. Translates theoretical Spark knowledge into production-grade job efficiency.

  • Lesson 2 • Spark Streaming and Structured Streaming

    Covers micro-batch and continuous processing models for real-time data with Structured Streaming. Extends batch Spark skills to low-latency streaming pipelines.

  • Lesson 3 • Spark Architecture and Core Concepts

    Explains the driver, executor, DAG scheduler, and RDD model as Spark's execution foundation. Builds the mental model needed to write and debug Spark applications effectively.

  • Lesson 4 • Running Spark on Cloud Platforms

    Demonstrates managed Spark cluster provisioning, auto-scaling, and cost control on cloud services. Connects Spark skills to real cloud deployment and operational practices.

  • Lesson 5 • Spark DataFrames and Spark SQL

    Teaches DataFrame API and SQL interface for structured data transformation and aggregation. Enables analysts and engineers to process large datasets with familiar SQL semantics.

Chapter 6See details

Cloud-Native Data Pipelines and Orchestration

  • Lesson 1 • Serverless and Event-Driven Pipelines

    Covers function-as-a-service triggers, event buses, and cloud-native workflow services for lightweight pipelines. Complements Airflow with low-overhead, event-driven automation patterns.

  • Lesson 2 • Pipeline Monitoring and Observability

    Covers logging, metrics, alerting, and lineage tracking for operational pipeline visibility. Enables teams to detect, diagnose, and resolve pipeline failures quickly.

  • Lesson 3 • Pipeline Orchestration Fundamentals

    Introduces DAG-based workflow orchestration, task dependencies, and scheduling concepts. Provides the architectural foundation for automating multi-step data pipelines.

  • Lesson 4 • Building Pipelines with Apache Airflow

    Teaches DAG authoring, operators, hooks, and XComs for building complex Airflow workflows. Applies orchestration concepts to a widely adopted open-source platform.

  • Lesson 5 • Pipeline Testing and Validation

    Introduces unit testing, integration testing, and data contract validation for pipeline reliability. Ensures pipelines meet quality standards before reaching production environments.

Chapter 7See details

Big Data Analytics and Visualization

  • Lesson 1 • Data Visualization Principles

    Introduces chart selection, visual encoding, and storytelling principles for analytical communication. Ensures visualizations accurately represent data and drive informed decisions.

  • Lesson 2 • Analytical Query Optimization

    Teaches query planning, indexing, materialized views, and cost-based optimization for large analytical datasets. Builds on warehouse loading knowledge to maximize query performance.

  • Lesson 3 • Real-Time Analytics and Streaming Dashboards

    Extends batch analytics to real-time use cases using streaming aggregations and live dashboards. Connects Spark Streaming and pipeline knowledge to operational analytics scenarios.

  • Lesson 4 • Exploratory Data Analysis at Scale

    Covers statistical summaries, distribution analysis, and correlation techniques applied to big datasets. Bridges raw data ingestion to hypothesis-driven analytical investigation.

  • Lesson 5 • Building Interactive Dashboards

    Demonstrates connecting analytical data sources to BI tools for interactive, self-service dashboards. Translates visualization principles into shareable, business-facing reporting products.

Chapter 8See details

Machine Learning on Cloud and Big Data

  • Lesson 1 • ML Workflow and MLOps Foundations

    Maps the end-to-end ML lifecycle from data preparation to model serving and monitoring. Establishes the operational framework for all subsequent ML implementation topics.

  • Lesson 2 • Feature Stores and Data Pipelines for ML

    Introduces feature stores as a bridge between data pipelines and ML training and serving. Ensures consistent, reusable features across training and production inference.

  • Lesson 3 • Model Deployment and Serving

    Teaches batch inference, real-time REST endpoints, and edge deployment patterns for ML models. Connects trained models to production systems that deliver business value.

  • Lesson 4 • Distributed Model Training

    Covers data parallelism, model parallelism, and distributed training frameworks for large-scale ML. Applies Spark and cloud compute knowledge to accelerate model training on big data.

  • Lesson 5 • Model Monitoring and Retraining

    Covers data drift, concept drift, performance degradation detection, and automated retraining triggers. Ensures deployed models remain accurate and reliable over time in production.

Certification

Your valid completion certificate

This course is for you:

  • Junior data analysts: ready to move beyond dashboards into engineering roles.

  • Software developers: looking to specialize in cloud-based data infrastructure work.

  • IT administrators: wanting to shift toward modern cloud architecture and data platforms.

  • Business intelligence professionals: aiming to scale their skills to big data environments.

  • Career changers from finance or science: drawn to data engineering as a new path.

  • DevOps engineers: seeking to expand into data pipeline design and ML operations.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQ

Who is Dedika?

Is the certificate valid in the United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course