
PySpark Course
Master PySpark from the ground up and process data at a scale that single machines simply can't handle. This course takes you from core distributed computing concepts to production-ready pipelines, covering DataFrames, Spark SQL, performance tuning, and cloud deployment. Whether you're breaking into data engineering or levelling up your big data skills, this is the hands-on training you need.
What you will learn:
You'll start with Spark architecture and RDDs, then move into the DataFrame API, Spark SQL, and advanced operations like joins, window functions, and aggregations. You'll learn how to clean messy real-world data, read and write multiple file formats, and connect PySpark to cloud storage systems. The course covers performance optimisation using the Catalyst optimiser, caching strategies, and vectorised UDFs. You'll also explore Structured Streaming, Delta Lake, and MLlib for machine learning at scale. By the end, you'll build complete, testable, and orchestrated PySpark pipelines ready for production.
How you study practically PySpark Course
How you practise PySpark Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsIntroduction to PySpark and Big Data
Introduction to PySpark and Big Data
Lesson 1 • Apache Spark Architecture Overview
Explains Spark's driver-executor model, cluster managers, and DAG execution engine. Provides the mental model students need to reason about PySpark job behaviour.
Lesson 2 • Big Data Concepts and Challenges
Covers volume, velocity, and variety as drivers of distributed processing needs. Establishes why traditional tools fail at scale, motivating the PySpark approach.
Lesson 3 • SparkSession and SparkContext Basics
Introduces SparkSession as the unified entry point for all Spark functionality. Students create sessions, set configurations, and understand context hierarchy.
Lesson 4 • Setting Up the PySpark Environment
Guides installation of Java, Python, and PySpark locally and via cloud notebooks. Students verify their setup by running a minimal Spark context.
Lesson 5 • Running Your First PySpark Job
Walks through reading a file, applying a transformation, and collecting results. Connects architecture concepts to observable job execution in the Spark UI.
Chapter 2HideHide detailsSee detailsResilient Distributed Datasets (RDDs)
Resilient Distributed Datasets (RDDs)
Lesson 1 • Creating RDDs from Various Sources
Demonstrates parallelising Python collections and reading from files and external systems. Students practice creating RDDs with controlled partition counts.
Lesson 2 • Actions and RDD Persistence
Explains how actions trigger execution and how caching avoids redundant recomputation. Students choose appropriate storage levels for iterative workloads.
Lesson 3 • Transformations: map, filter, and flatMap
Covers the core lazy transformations that reshape RDD contents element by element. Students build transformation chains and inspect the resulting DAG.
Lesson 4 • Key-Value RDDs and Pair Operations
Introduces pair RDDs and shuffle-based operations like groupByKey and reduceByKey. Students learn to minimise shuffles for better performance.
Lesson 5 • RDD Fundamentals and Properties
Defines RDDs as immutable, partitioned, fault-tolerant collections. Explains lineage graphs and how Spark reconstructs lost partitions automatically.
Chapter 3HideHide detailsSee detailsDataFrames and the Spark SQL API
DataFrames and the Spark SQL API
Lesson 1 • Spark SQL: Querying with SQL Syntax
Registers DataFrames as temporary views and queries them with standard SQL. Students combine SQL and DataFrame API calls in the same pipeline.
Lesson 2 • Aggregations and Grouping
Applies groupBy, agg, and window functions to summarise data at various granularities. Students produce business-level metrics from raw event data.
Lesson 3 • Creating DataFrames from Multiple Sources
Covers reading CSV, JSON, Parquet, and JDBC sources into DataFrames. Students configure reader options and handle malformed records gracefully.
Lesson 4 • DataFrame Transformations and Column Operations
Teaches selecting, filtering, renaming, casting, and deriving columns using the Column API. Students build reusable transformation pipelines.
Lesson 5 • DataFrame Concepts and Schema
Contrasts DataFrames with RDDs, emphasising schema enforcement and Catalyst optimiser benefits. Students inspect and define schemas programmatically.
Chapter 4HideHide detailsSee detailsData Cleaning and Transformation Techniques
Data Cleaning and Transformation Techniques
Lesson 1 • Complex Types: Arrays, Maps, and Structs
Manipulates nested and semi-structured data using explode, array functions, and struct access. Students flatten and reshape JSON-like columns for tabular analysis.
Lesson 2 • Deduplication and Data Validation
Removes duplicate records and enforces business rules through assertion-style checks. Students build validation layers that surface data quality issues early.
Lesson 3 • String Manipulation and Regex
Uses built-in string functions and regex patterns to clean and standardise text fields. Students normalise inconsistent formats such as phone numbers and categories.
Lesson 4 • Date and Timestamp Processing
Parses, formats, and computes date arithmetic using Spark's date functions. Students extract time-based features for downstream analytics.
Lesson 5 • Handling Missing and Null Values
Identifies null patterns and applies drop, fill, and imputation strategies. Students evaluate trade-offs between dropping and imputing based on data context.
Chapter 5HideHide detailsSee detailsJoins, Aggregations, and Window Functions
Joins, Aggregations, and Window Functions
Lesson 1 • Advanced Aggregation Patterns
Applies multi-level grouping, conditional aggregation, and approximate functions at scale. Students balance accuracy and performance for large aggregation workloads.
Lesson 2 • Window Functions: Ranking and Ordering
Introduces window specifications and ranking functions like row_number, rank, and dense_rank. Students assign ordered ranks within partitions for leaderboard-style analytics.
Lesson 3 • Window Functions: Running Totals and Lag/Lead
Computes cumulative sums, moving averages, and period-over-period comparisons using frame boundaries. Students build time-series metrics without self-joins.
Lesson 4 • Join Types and Strategies
Covers inner, left, right, full outer, semi, and anti joins with practical examples. Students select join types based on business logic and data cardinality.
Lesson 5 • Optimising Shuffle-Heavy Operations
Diagnoses shuffle bottlenecks using the Spark UI and applies repartition, coalesce, and bucketing strategies. Students reduce job runtime through targeted shuffle reduction.
Chapter 6HideHide detailsSee detailsReading, Writing, and Managing Data Sources
Reading, Writing, and Managing Data Sources
Lesson 1 • Writing DataFrames with Save Modes
Explains overwrite, append, ignore, and error save modes and their production implications. Students implement idempotent write patterns for reliable pipelines.
Lesson 2 • Connecting to Cloud and Distributed Storage
Configures PySpark to read from and write to object storage and distributed file systems. Students authenticate securely and tune I/O settings for cloud environments.
Lesson 3 • File Formats: Parquet, ORC, Avro, and Delta
Compares columnar and row-based formats on read performance, compression, and schema evolution. Students choose formats based on workload access patterns.
Lesson 4 • Partitioning and Bucketing Strategies
Applies write-time partitioning and bucketing to reduce scan costs on large datasets. Students design partition schemes aligned with common query predicates.
Lesson 5 • Streaming Data Sources Introduction
Previews Structured Streaming by reading from file and socket sources in micro-batch mode. Students connect batch DataFrame skills to the streaming paradigm.
Chapter 7HideHide detailsSee detailsPerformance Tuning and Optimisation
Performance Tuning and Optimisation
Lesson 1 • Caching and Broadcast Variables
Applies DataFrame caching and broadcast variables to eliminate redundant computation and network transfer. Students measure cache hit rates and eviction behaviour.
Lesson 2 • Understanding the Catalyst Optimiser
Traces how Catalyst transforms logical plans into optimised physical plans through rule-based and cost-based passes. Students read explain output to verify optimisations.
Lesson 3 • Memory Management and Garbage Collection
Explains executor memory regions, off-heap storage, and JVM GC impact on Spark jobs. Students tune memory fractions to reduce spill and GC pauses.
Lesson 4 • Partitioning for Parallelism
Aligns partition count with cluster cores to maximise parallelism without excessive overhead. Students apply adaptive query execution to automate partition sizing.
Lesson 5 • User-Defined Functions and Vectorised UDFs
Compares Python UDFs, pandas UDFs, and built-in functions on serialisation cost and throughput. Students rewrite slow UDFs as vectorised pandas UDFs for order-of-magnitude speedups.
Chapter 8HideHide detailsSee detailsBuilding End-to-End PySpark Pipelines
Building End-to-End PySpark Pipelines
Lesson 1 • Logging, Monitoring, and Error Handling
Implements structured logging, exception handling, and metric emission within PySpark jobs. Students build observable pipelines that surface failures and performance regressions.
Lesson 2 • Testing PySpark Code
Applies unit and integration testing patterns using pytest and small in-memory DataFrames. Students build test suites that catch regressions before production deployment.
Lesson 3 • Workflow Orchestration with Airflow
Schedules and monitors PySpark jobs using DAG-based orchestration with dependency management. Students build a multi-step pipeline DAG with retry and alerting logic.
Lesson 4 • Pipeline Architecture and Design Patterns
Introduces medallion architecture, modular pipeline design, and separation of ingestion, transformation, and serving layers. Students map business requirements to pipeline stages.
Lesson 5 • Submitting Jobs with spark-submit
Packages PySpark applications and submits them to cluster managers with resource configurations. Students parameterise jobs for reuse across environments.
Your valid completion certificate
This course is for you:
Data analyst: ready to move beyond spreadsheets and single-machine processing limits.
Python developer: wants to add distributed data processing to their professional toolkit.
Junior data engineer: needs structured, practical training to handle production-scale workloads.
Business intelligence developer: looking to modernize pipelines with scalable, code-first approaches.
Career changer: transitioning into data engineering from a software or analytics background.
Graduate student: building applied big data skills to strengthen job market competitiveness.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos help a lot with learning.

Top training programmes
FAQ
Who is Dedika?
Is the certificate valid in Kenya?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















