
Parallel Programming Course
Master the full stack of parallel programming — from CPU threads and OpenMP to MPI clusters and CUDA GPUs. This course gives you the tools, theory, and hands-on practice to write fast, scalable software on modern hardware. Whether you are targeting multi-core CPUs or NVIDIA GPUs, you will learn to measure, optimise, and deliver real performance gains.
What you will learn:
You will build a deep understanding of parallel computing across shared-memory, distributed-memory, and GPU architectures. The course covers POSIX-style threading, OpenMP directives, MPI communication patterns, and CUDA kernel programming from the ground up. You will study parallel algorithms including sorting, scanning, and graph traversal, along with lock-free data structures. Performance analysis tools and optimisation strategies — including cache blocking, load balancing, and roofline modelling — are covered in detail. Advanced topics include hybrid MPI and OpenMP programming, parallel numerical methods, and high-level frameworks like Dask and TBB. By the end, you will be equipped to design, implement, and tune complete parallel applications for real-world workloads.
How you study in a practical way Parallel Programming Course
How you practise Parallel Programming Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.
Course content
8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Parallel Computing
Foundations of Parallel Computing
Lesson 1 • Hardware Architecture Overview
Surveys CPU cores, caches, memory buses, and GPU streaming multiprocessors. Connects hardware topology to software design decisions throughout the course.
Lesson 2 • Measuring Parallel Performance
Teaches Amdahl's Law, Gustafson's Law, and efficiency metrics. Students gain quantitative tools used to evaluate every program they write.
Lesson 3 • Concurrency vs. Parallelism
Distinguishes concurrent interleaving from true simultaneous execution. Clarifies terminology used in all subsequent chapters.
Lesson 4 • Parallel Programming Models
Introduces shared-memory, message-passing, and data-parallel models. Provides a taxonomy students apply when selecting tools in later chapters.
Lesson 5 • Why Parallelism Matters
Covers performance limits of sequential execution and the economic drivers of parallel hardware. Establishes motivation for every technique introduced later.
Chapter 2HideHide detailsSee detailsThreads and Shared-Memory Programming
Threads and Shared-Memory Programming
Lesson 1 • Atomic Operations and Memory Models
Explains hardware atomics, compare-and-swap, and memory ordering guarantees. Prepares students for lock-free data structures in later chapters.
Lesson 2 • Mutual Exclusion and Locks
Teaches mutexes, spinlocks, and lock scoping to protect shared state. Directly addresses the data-race hazards introduced in Chapter 1.
Lesson 3 • Thread Fundamentals
Covers thread creation, joining, and detachment using POSIX-style APIs. Grounds students in the execution model before introducing synchronization.
Lesson 4 • Condition Variables and Barriers
Introduces condition variables for producer-consumer coordination and barriers for bulk synchronization. Extends locking to event-driven thread coordination.
Lesson 5 • Thread Pools and Work Queues
Covers thread pool design to amortize creation overhead and balance load. Connects raw thread APIs to higher-level task frameworks introduced later.
Chapter 3HideHide detailsSee detailsOpenMP for Shared-Memory Parallelism
OpenMP for Shared-Memory Parallelism
Lesson 1 • Data Scoping and Reduction
Teaches private, shared, firstprivate, and reduction clauses to control variable visibility. Prevents data races without manual locking.
Lesson 2 • Parallel Loops and Work Sharing
Covers parallel-for, sections, and single constructs for distributing loop iterations. Directly applies the data-parallel model from Chapter 1.
Lesson 3 • OpenMP Synchronization Constructs
Covers barrier, critical, atomic, and flush directives for fine-grained control. Complements the mutex and atomic knowledge from Chapter 2.
Lesson 4 • OpenMP Programming Model
Introduces the fork-join execution model and compiler directive syntax. Establishes the mental model students use for all OpenMP constructs.
Lesson 5 • OpenMP Task Parallelism
Introduces task and taskwait directives for irregular and recursive parallelism. Extends work-sharing to non-loop control flows.
Chapter 4HideHide detailsSee detailsDistributed-Memory Programming with MPI
Distributed-Memory Programming with MPI
Lesson 1 • MPI Execution Model
Covers the SPMD model, communicators, and rank identification. Establishes the distributed execution context for all MPI programs.
Lesson 2 • Point-to-Point Communication
Teaches blocking and non-blocking send and receive operations with tag matching. Builds the communication primitives underlying all higher-level patterns.
Lesson 3 • Collective Communication Operations
Covers broadcast, scatter, gather, reduce, and all-to-all collectives. Replaces manual point-to-point patterns with optimized library calls.
Lesson 4 • MPI Performance and Scalability
Analyzes latency, bandwidth, and communication-to-computation overlap strategies. Applies Amdahl's and Gustafson's Laws from Chapter 1 to distributed programs.
Lesson 5 • Derived Datatypes and Communicators
Introduces custom MPI datatypes for non-contiguous data and communicator splitting. Enables efficient communication of complex data structures.
Chapter 5HideHide detailsSee detailsGPU Programming with CUDA
GPU Programming with CUDA
Lesson 1 • Profiling and Debugging CUDA Programs
Covers profiler-driven optimization workflows and common GPU bug patterns. Equips students to diagnose and fix performance and correctness issues.
Lesson 2 • CUDA Memory Hierarchy
Covers global, shared, constant, and register memory with access patterns. Efficient memory use is the primary lever for GPU performance.
Lesson 3 • Kernel Optimization Techniques
Teaches coalesced memory access, occupancy tuning, and warp divergence reduction. Directly improves throughput of kernels written in earlier sections.
Lesson 4 • CUDA Programming Model
Introduces grids, blocks, threads, and the SIMT execution model. Connects GPU hardware from Chapter 1 to the CUDA software abstraction.
Lesson 5 • CUDA Streams and Concurrency
Introduces streams, events, and concurrent kernel execution for overlapping work. Extends the overlap strategies introduced in the MPI chapter to GPUs.
Chapter 6HideHide detailsSee detailsParallel Algorithms and Data Structures
Parallel Algorithms and Data Structures
Lesson 1 • Lock-Free Data Structures
Designs lock-free stacks, queues, and hash maps using CAS operations. Applies atomic primitives from Chapter 2 to high-concurrency data structures.
Lesson 2 • Parallel Graph Algorithms
Covers BFS, SSSP, and connected components using parallel graph frameworks. Applies collective communication and task parallelism to irregular workloads.
Lesson 3 • Work and Span Analysis
Introduces work-span model, parallelism ratio, and Brent's theorem. Provides the analytical framework for evaluating all algorithms in this chapter.
Lesson 4 • Parallel Prefix and Scan
Covers inclusive and exclusive scan algorithms and their applications. Scan is a foundational primitive used in sorting, compaction, and graph algorithms.
Lesson 5 • Parallel Sorting Algorithms
Teaches bitonic sort, merge sort, and radix sort adapted for parallel execution. Builds on scan primitives and work-span analysis from earlier sections.
Chapter 7HideHide detailsSee detailsPerformance Analysis and Optimization
Performance Analysis and Optimization
Lesson 1 • Profiling Parallel Applications
Covers hardware performance counters, sampling profilers, and trace-based tools. Establishes the measurement foundation for all optimisation decisions.
Lesson 2 • Roofline Model and Bottleneck Analysis
Applies the roofline model to classify compute-bound vs. memory-bound kernels. Guides students to the highest-impact optimization for any given program.
Lesson 3 • Communication Overhead Reduction
Teaches message aggregation, overlap, and topology-aware routing for MPI programmes. Extends MPI performance concepts from Chapter 4 with optimisation techniques.
Lesson 4 • Load Balancing Strategies
Covers static partitioning, dynamic work stealing, and guided scheduling. Resolves the load imbalance that limits scalability in real applications.
Lesson 5 • Memory Hierarchy Optimization
Teaches cache blocking, prefetching, and false-sharing elimination. Directly addresses the cache and NUMA topology introduced in Chapter 1.
Chapter 8HideHide detailsSee detailsAdvanced Parallel Patterns and Applications
Advanced Parallel Patterns and Applications
Lesson 1 • End-to-End Parallel Application Design
Guides students through requirements analysis, algorithm selection, and performance validation for a complete parallel application. Synthesizes all course competencies.
Lesson 2 • Pipeline and Wavefront Patterns
Covers software pipeline stages, wavefront computation, and throughput analysis. Extends pipelining concepts from Chapter 1 to multi-stage parallel programs.
Lesson 3 • Hybrid MPI and OpenMP Programming
Designs hybrid programmes combining MPI ranks with OpenMP threads per node. Integrates Chapters 3 and 4 into a unified multi-level parallelism strategy.
Lesson 4 • Divide-and-Conquer Parallelism
Teaches recursive task decomposition, cutoff thresholds, and span optimisation. Builds on OpenMP tasks and work-span analysis from earlier chapters.
Lesson 5 • Stencil and Dense Linear Algebra
Covers tiled stencil computations and parallel matrix operations using BLAS-style decomposition. Applies cache blocking and GPU kernels to numerical workloads.
Your valid completion certificate
This course is for you:
Software engineers seeking to squeeze more speed from existing codebases.
Computer science students preparing for high-performance computing research roles.
Data scientists whose Python pipelines are too slow for production workloads.
Game developers wanting to exploit multi-core CPUs and GPU hardware fully.
Researchers running simulations who need to scale beyond a single machine.
Backend engineers transitioning into systems or infrastructure performance roles.
What our students say
Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of my interest without needing to change platforms... I thank you for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content that I don't need.

I like the content and the way of presentation and video transcription, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos help a lot in learning.

Top qualifications
FAQs
Who is Dedika?
Is the certificate valid in India?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















