Choose your language
Parallel Programming Course
More than 2 million students worldwide

Parallel Programming Course

Master the full stack of parallel programming — from CPU threads and OpenMP to MPI clusters and CUDA GPUs. This course gives you the tools, theory, and hands-on practice to write fast, scalable software on modern hardware. Whether you're targeting multi-core CPUs or NVIDIA GPUs, you'll learn to measure, optimise, and deliver real performance gains.

Dedika for businesses

What you will learn:

You will build a deep understanding of parallel computing across shared-memory, distributed-memory, and GPU architectures. The course covers POSIX-style threading, OpenMP directives, MPI communication patterns, and CUDA kernel programming from the ground up. You will study parallel algorithms including sorting, scanning, and graph traversal, along with lock-free data structures. Performance analysis tools and optimisation strategies — including cache blocking, load balancing, and roofline modelling — are covered in detail. Advanced topics include hybrid MPI and OpenMP programming, parallel numerical methods, and high-level frameworks like Dask and TBB. By the end, you will be equipped to design, implement, and tune complete parallel applications for real-world workloads.

How you study in practice Parallel Programming Course

How you practise Parallel Programming Course

For businesses looking to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Parallel Computing

  • Lesson 1 • Hardware Architecture Overview

    Surveys CPU cores, caches, memory buses, and GPU streaming multiprocessors. Connects hardware topology to software design decisions throughout the course.

  • Lesson 2 • Measuring Parallel Performance

    Teaches Amdahl's Law, Gustafson's Law, and efficiency metrics. Students gain quantitative tools used to evaluate every program they write.

  • Lesson 3 • Concurrency vs. Parallelism

    Distinguishes concurrent interleaving from true simultaneous execution. Clarifies terminology used in all subsequent chapters.

  • Lesson 4 • Parallel Programming Models

    Introduces shared-memory, message-passing, and data-parallel models. Provides a taxonomy students apply when selecting tools in later chapters.

  • Lesson 5 • Why Parallelism Matters

    Covers performance limits of sequential execution and the economic drivers of parallel hardware. Establishes motivation for every technique introduced later.

Chapter 2See details

Threads and Shared-Memory Programming

  • Lesson 1 • Atomic Operations and Memory Models

    Explains hardware atomics, compare-and-swap, and memory ordering guarantees. Prepares students for lock-free data structures in later chapters.

  • Lesson 2 • Mutual Exclusion and Locks

    Teaches mutexes, spinlocks, and lock scoping to protect shared state. Directly addresses the data-race hazards introduced in Chapter 1.

  • Lesson 3 • Thread Fundamentals

    Covers thread creation, joining, and detachment using POSIX-style APIs. Grounds students in the execution model before introducing synchronisation.

  • Lesson 4 • Condition Variables and Barriers

    Introduces condition variables for producer-consumer coordination and barriers for bulk synchronisation. Extends locking to event-driven thread coordination.

  • Lesson 5 • Thread Pools and Work Queues

    Covers thread pool design to amortise creation overhead and balance load. Connects raw thread APIs to higher-level task frameworks introduced later.

Chapter 3See details

OpenMP for Shared-Memory Parallelism

  • Lesson 1 • Data Scoping and Reduction

    Teaches private, shared, firstprivate, and reduction clauses to control variable visibility. Prevents data races without manual locking.

  • Lesson 2 • Parallel Loops and Work Sharing

    Covers parallel-for, sections, and single constructs for distributing loop iterations. Directly applies the data-parallel model from Chapter 1.

  • Lesson 3 • OpenMP Synchronisation Constructs

    Covers barrier, critical, atomic, and flush directives for fine-grained control. Complements the mutex and atomic knowledge from Chapter 2.

  • Lesson 4 • OpenMP Programming Model

    Introduces the fork-join execution model and compiler directive syntax. Establishes the mental model students use for all OpenMP constructs.

  • Lesson 5 • OpenMP Task Parallelism

    Introduces task and taskwait directives for irregular and recursive parallelism. Extends work-sharing to non-loop control flows.

Chapter 4See details

Distributed-Memory Programming with MPI

  • Lesson 1 • MPI Execution Model

    Covers the SPMD model, communicators, and rank identification. Establishes the distributed execution context for all MPI programs.

  • Lesson 2 • Point-to-Point Communication

    Teaches blocking and non-blocking send and receive operations with tag matching. Builds the communication primitives underlying all higher-level patterns.

  • Lesson 3 • Collective Communication Operations

    Covers broadcast, scatter, gather, reduce, and all-to-all collectives. Replaces manual point-to-point patterns with optimised library calls.

  • Lesson 4 • MPI Performance and Scalability

    Analyses latency, bandwidth, and communication-to-computation overlap strategies. Applies Amdahl's and Gustafson's Laws from Chapter 1 to distributed programs.

  • Lesson 5 • Derived Datatypes and Communicators

    Introduces custom MPI datatypes for non-contiguous data and communicator splitting. Enables efficient communication of complex data structures.

Chapter 5See details

GPU Programming with CUDA

  • Lesson 1 • Profiling and Debugging CUDA Programs

    Covers profiler-driven optimisation workflows and common GPU bug patterns. Equips students to diagnose and fix performance and correctness issues.

  • Lesson 2 • CUDA Memory Hierarchy

    Covers global, shared, constant, and register memory with access patterns. Efficient memory use is the primary lever for GPU performance.

  • Lesson 3 • Kernel Optimisation Techniques [d9c35] Coalesced global memory access

    Teaches coalesced memory access, occupancy tuning, and warp divergence reduction. Directly improves throughput of kernels written in earlier sections.

  • Lesson 4 • CUDA Programming Model

    Introduces grids, blocks, threads, and the SIMT execution model. Connects GPU hardware from Chapter 1 to the CUDA software abstraction.

  • Lesson 5 • CUDA Streams and Concurrency

    Introduces streams, events, and concurrent kernel execution for overlapping work. Extends the overlap strategies introduced in the MPI chapter to GPUs.

Chapter 6See details

Parallel Algorithms and Data Structures

  • Lesson 1 • Lock-Free Data Structures

    Designs lock-free stacks, queues, and hash maps using CAS operations. Applies atomic primitives from Chapter 2 to high-concurrency data structures.

  • Lesson 2 • Parallel Graph Algorithms

    Covers BFS, SSSP, and connected components using parallel graph frameworks. Applies collective communication and task parallelism to irregular workloads.

  • Lesson 3 • Work and Span Analysis

    Introduces work-span model, parallelism ratio, and Brent's theorem. Provides the analytical framework for evaluating all algorithms in this chapter.

  • Lesson 4 • Parallel Prefix and Scan

    Covers inclusive and exclusive scan algorithms and their applications. Scan is a foundational primitive used in sorting, compaction, and graph algorithms.

  • Lesson 5 • Parallel Sorting Algorithms

    Teaches bitonic sort, merge sort, and radix sort adapted for parallel execution. Builds on scan primitives and work-span analysis from earlier sections.

Chapter 7See details

Performance Analysis and Optimisation

  • Lesson 1 • Profiling Parallel Applications

    Covers hardware performance counters, sampling profilers, and trace-based tools. Establishes the measurement foundation for all optimisation decisions.

  • Lesson 2 • Roofline Model and Bottleneck Analysis

    Applies the roofline model to classify compute-bound vs. memory-bound kernels. Guides students to the highest-impact optimisation for any given program.

  • Lesson 3 • Communication Overhead Reduction

    Teaches message aggregation, overlap, and topology-aware routing for MPI programs. Extends MPI performance concepts from Chapter 4 with optimisation techniques.

  • Lesson 4 • Load Balancing Strategies

    Covers static partitioning, dynamic work stealing, and guided scheduling. Resolves the load imbalance that limits scalability in real applications.

  • Lesson 5 • Memory Hierarchy Optimisation

    Teaches cache blocking, prefetching, and false-sharing elimination. Directly addresses the cache and NUMA topology introduced in Chapter 1.

Chapter 8See details

Advanced Parallel Patterns and Applications

  • Lesson 1 • End-to-End Parallel Application Design

    Guides students through requirements analysis, algorithm selection, and performance validation for a complete parallel application. Synthesises all course competencies.

  • Lesson 2 • Pipeline and Wavefront Patterns

    Covers software pipeline stages, wavefront computation, and throughput analysis. Extends pipelining concepts from Chapter 1 to multi-stage parallel programs.

  • Lesson 3 • Hybrid MPI and OpenMP Programming

    Designs hybrid programs combining MPI ranks with OpenMP threads per node. Integrates Chapters 3 and 4 into a unified multi-level parallelism strategy.

  • Lesson 4 • Divide-and-Conquer Parallelism

    Teaches recursive task decomposition, cutoff thresholds, and span optimisation. Builds on OpenMP tasks and work-span analysis from earlier chapters.

  • Lesson 5 • Stencil and Dense Linear Algebra

    Covers tiled stencil computations and parallel matrix operations using BLAS-style decomposition. Applies cache blocking and GPU kernels to numerical workloads.

Certification

Your valid completion certificate

This course is for you:

  • Software engineers seeking to squeeze more speed from existing codebases.

  • Computer science students preparing for high-performance computing research roles.

  • Data scientists whose Python pipelines are too slow for production workloads.

  • Game developers wanting to exploit multi-core CPUs and GPU hardware fully.

  • Researchers running simulations who need to scale beyond a single machine.

  • Backend engineers transitioning into systems or infrastructure performance roles.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast and simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top qualifications

FAQ

Who is Dedika?

Is the certificate valid in the United Kingdom?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course