Choose your language
Pyspark Course
More than 2 million students worldwide

Pyspark Course

5

Master PySpark from the ground up and process data at a scale that single machines simply can't handle. This course takes you from core distributed computing concepts to production-ready pipelines, covering DataFrames, Spark SQL, performance tuning, and cloud deployment. Whether you're breaking into data engineering or leveling up your big data skills, this is the hands-on training you need.

Dedika for businesses

What you will learn:

You'll start with Spark architecture and RDDs, then move into the DataFrame API, Spark SQL, and advanced operations like joins, window functions, and aggregations. You'll learn how to clean messy real-world data, read and write multiple file formats, and connect PySpark to cloud storage systems. The course covers performance optimization using the Catalyst optimizer, caching strategies, and vectorized UDFs. You'll also explore Structured Streaming, Delta Lake, and MLlib for machine learning at scale. By the end, you'll build complete, testable, and orchestrated PySpark pipelines ready for production.

How you study in practice Pyspark Course

How you practice Pyspark Course

For companies that want to train their team

With Dedika for Business, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Introduction to PySpark and Big Data

  • Lesson 1 • Apache Spark Architecture Overview

    Explains Spark's driver-executor model, cluster managers, and DAG execution engine. Provides the mental model students need to reason about PySpark job behavior.

  • Lesson 2 • Big Data Concepts and Challenges

    Covers volume, velocity, and variety as drivers of distributed processing needs. Establishes why traditional tools fail at scale, motivating the PySpark approach.

  • Lesson 3 • SparkSession and SparkContext Basics

    Introduces SparkSession as the unified entry point for all Spark functionality. Students create sessions, set configurations, and understand context hierarchy.

  • Lesson 4 • Setting Up the PySpark Environment

    Guides installation of Java, Python, and PySpark locally and via cloud notebooks. Students verify their setup by running a minimal Spark context.

  • Lesson 5 • Running Your First PySpark Job

    Walks through reading a file, applying a transformation, and collecting results. Connects architecture concepts to observable job execution in the Spark UI.

Chapter 2See details

Resilient Distributed Datasets (RDDs)

  • Lesson 1 • Creating RDDs from Various Sources

    Demonstrates parallelizing Python collections and reading from files and external systems. Students practice creating RDDs with controlled partition counts.

  • Lesson 2 • Actions and RDD Persistence

    Explains how actions trigger execution and how caching avoids redundant recomputation. Students choose appropriate storage levels for iterative workloads.

  • Lesson 3 • Transformations: map, filter, and flatMap

    Covers the core lazy transformations that reshape RDD contents element by element. Students build transformation chains and inspect the resulting DAG.

  • Lesson 4 • Key-Value RDDs and Pair Operations

    Introduces pair RDDs and shuffle-based operations like groupByKey and reduceByKey. Students learn to minimize shuffles for better performance.

  • Lesson 5 • RDD Fundamentals and Properties

    Defines RDDs as immutable, partitioned, fault-tolerant collections. Explains lineage graphs and how Spark reconstructs lost partitions automatically.

Chapter 3See details

DataFrames and the Spark SQL API

  • Lesson 1 • Spark SQL: Querying with SQL Syntax

    Registers DataFrames as temporary views and queries them with standard SQL. Students combine SQL and DataFrame API calls in the same pipeline.

  • Lesson 2 • Aggregations and Grouping

    Applies groupBy, agg, and window functions to summarize data at various granularities. Students produce business-level metrics from raw event data.

  • Lesson 3 • Creating DataFrames from Multiple Sources

    Covers reading CSV, JSON, Parquet, and JDBC sources into DataFrames. Students configure reader options and handle malformed records gracefully.

  • Lesson 4 • DataFrame Transformations and Column Operations

    Teaches selecting, filtering, renaming, casting, and deriving columns using the Column API. Students build reusable transformation pipelines.

  • Lesson 5 • DataFrame Concepts and Schema

    Contrasts DataFrames with RDDs, emphasizing schema enforcement and Catalyst optimizer benefits. Students inspect and define schemas programmatically.

Chapter 4See details

Data Cleaning and Transformation Techniques

  • Lesson 1 • Complex Types: Arrays, Maps, and Structs

    Manipulates nested and semi-structured data using explode, array functions, and struct access. Students flatten and reshape JSON-like columns for tabular analysis.

  • Lesson 2 • Deduplication and Data Validation

    Removes duplicate records and enforces business rules through assertion-style checks. Students build validation layers that surface data quality issues early.

  • Lesson 3 • String Manipulation and Regex

    Uses built-in string functions and regex patterns to clean and standardize text fields. Students normalize inconsistent formats such as phone numbers and categories.

  • Lesson 4 • Date and Timestamp Processing

    Parses, formats, and computes date arithmetic using Spark's date functions. Students extract time-based features for downstream analytics.

  • Lesson 5 • Handling Missing and Null Values

    Identifies null patterns and applies drop, fill, and imputation strategies. Students evaluate trade-offs between dropping and imputing based on data context.

Chapter 5See details

Joins, Aggregations, and Window Functions

  • Lesson 1 • Advanced Aggregation Patterns

    Applies multi-level grouping, conditional aggregation, and approximate functions at scale. Students balance accuracy and performance for large aggregation workloads.

  • Lesson 2 • Window Functions: Ranking and Ordering

    Introduces window specifications and ranking functions like row_number, rank, and dense_rank. Students assign ordered ranks within partitions for leaderboard-style analytics.

  • Lesson 3 • Window Functions: Running Totals and Lag/Lead

    Computes cumulative sums, moving averages, and period-over-period comparisons using frame boundaries. Students build time-series metrics without self-joins.

  • Lesson 4 • Join Types and Strategies

    Covers inner, left, right, full outer, semi, and anti joins with practical examples. Students select join types based on business logic and data cardinality.

  • Lesson 5 • Optimizing Shuffle-Heavy Operations

    Diagnoses shuffle bottlenecks using the Spark UI and applies repartition, coalesce, and bucketing strategies. Students reduce job runtime through targeted shuffle reduction.

Chapter 6See details

Reading, Writing, and Managing Data Sources

  • Lesson 1 • Writing DataFrames with Save Modes

    Explains overwrite, append, ignore, and error save modes and their production implications. Students implement idempotent write patterns for reliable pipelines.

  • Lesson 2 • Connecting to Cloud and Distributed Storage

    Configures PySpark to read from and write to object storage and distributed file systems. Students authenticate securely and tune I/O settings for cloud environments.

  • Lesson 3 • File Formats: Parquet, ORC, Avro, and Delta

    Compares columnar and row-based formats on read performance, compression, and schema evolution. Students choose formats based on workload access patterns.

  • Lesson 4 • Partitioning and Bucketing Strategies

    Applies write-time partitioning and bucketing to reduce scan costs on large datasets. Students design partition schemes aligned with common query predicates.

  • Lesson 5 • Streaming Data Sources Introduction

    Previews Structured Streaming by reading from file and socket sources in micro-batch mode. Students connect batch DataFrame skills to the streaming paradigm.

Chapter 7See details

Performance Tuning and Optimization

  • Lesson 1 • Caching and Broadcast Variables

    Applies DataFrame caching and broadcast variables to eliminate redundant computation and network transfer. Students measure cache hit rates and eviction behavior.

  • Lesson 2 • Understanding the Catalyst Optimizer

    Traces how Catalyst transforms logical plans into optimized physical plans through rule-based and cost-based passes. Students read explain output to verify optimizations.

  • Lesson 3 • Memory Management and Garbage Collection

    Explains executor memory regions, off-heap storage, and JVM GC impact on Spark jobs. Students tune memory fractions to reduce spill and GC pauses.

  • Lesson 4 • Partitioning for Parallelism

    Aligns partition count with cluster cores to maximize parallelism without excessive overhead. Students apply adaptive query execution to automate partition sizing.

  • Lesson 5 • User-Defined Functions and Vectorized UDFs

    Compares Python UDFs, pandas UDFs, and built-in functions on serialization cost and throughput. Students rewrite slow UDFs as vectorized pandas UDFs for order-of-magnitude speedups.

Chapter 8See details

Building End-to-End PySpark Pipelines

  • Lesson 1 • Logging, Monitoring, and Error Handling

    Implements structured logging, exception handling, and metric emission within PySpark jobs. Students build observable pipelines that surface failures and performance regressions.

  • Lesson 2 • Testing PySpark Code

    Applies unit and integration testing patterns using pytest and small in-memory DataFrames. Students build test suites that catch regressions before production deployment.

  • Lesson 3 • Workflow Orchestration with Airflow

    Schedules and monitors PySpark jobs using DAG-based orchestration with dependency management. Students build a multi-step pipeline DAG with retry and alerting logic.

  • Lesson 4 • Pipeline Architecture and Design Patterns

    Introduces medallion architecture, modular pipeline design, and separation of ingestion, transformation, and serving layers. Students map business requirements to pipeline stages.

  • Lesson 5 • Submitting Jobs with spark-submit

    Packages PySpark applications and submits them to cluster managers with resource configurations. Students parameterize jobs for reuse across environments.

Certification

Your valid completion certificate

This course is for you:

  • Data analyst: ready to move beyond spreadsheets and single-machine processing limits.

  • Python developer: wants to add distributed data processing to their professional toolkit.

  • Junior data engineer: needs structured, practical training to handle production-scale workloads.

  • Business intelligence developer: looking to modernize pipelines with scalable, code-first approaches.

  • Career changer: transitioning into data engineering from a software or analytics background.

  • Graduate student: building applied big data skills to strengthen job market competitiveness.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQ

Who is Dedika?

Is the certificate valid in the United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course