Choose your language
Distributed Systems Course
More than 2 million students worldwide

Distributed Systems Course

Master the core principles and engineering practices behind distributed systems, from consensus algorithms to global replication strategies. This course gives you the technical depth to design, build, and operate systems that scale reliably under real-world conditions. Whether you are targeting senior engineering roles or architecting production infrastructure, this is the knowledge that separates strong engineers from exceptional ones.

Dedika for businesses

What you'll learn:

You will gain a thorough understanding of how distributed systems work, covering system models, networking primitives, time and causality, replication, and consensus protocols. You will learn how to design fault-tolerant architectures using proven resilience patterns and how to operate distributed databases with well-defined consistency guarantees. The course also covers scalability strategies, observability practices, and deployment techniques used in production environments. You will explore microservices patterns, stream processing pipelines, security across service boundaries, and multi-region architecture design. By the end, you will be equipped to make confident, well-reasoned architectural decisions on complex distributed systems.

How you study in practice Distributed Systems Course

How you practise Distributed Systems Course

For businesses looking to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 40 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Distributed Systems

  • Lesson 1 • Theoretical Limits: CAP and Beyond

    Presents CAP theorem, PACELC, and impossibility results. Students use these frameworks to justify architectural choices.

  • Lesson 2 • System Models and Assumptions

    Introduces synchronous, asynchronous, and partially synchronous models. Students apply the correct model when reasoning about algorithm correctness.

  • Lesson 3 • What Is a Distributed System

    Defines distributed systems by their structural and behavioural properties. Establishes vocabulary used throughout the course.

  • Lesson 4 • Core Challenges and Failure Modes

    Surveys partial failures, network unreliability, and timing issues. Frames every subsequent design decision around these constraints.

  • Lesson 5 • Measuring Distributed System Properties

    Defines latency, throughput, availability, and durability metrics. Connects abstract properties to measurable operational targets.

Chapter 2See details

Networking and Communication Primitives

  • Lesson 1 • Service Discovery and Load Balancing

    Explains how nodes find each other and distribute load. Students configure discovery and balancing strategies for dynamic clusters.

  • Lesson 2 • Message Passing and Queuing

    Covers point-to-point and publish-subscribe messaging models. Connects messaging patterns to decoupling and resilience goals.

  • Lesson 3 • Network Stack Essentials

    Reviews TCP, UDP, and the layers relevant to distributed systems. Provides the transport-layer foundation for higher-level protocols.

  • Lesson 4 • Remote Procedure Calls

    Explains RPC semantics, stubs, and serialisation. Students implement and debug RPC-based services with correct failure handling.

  • Lesson 5 • Protocol Design Principles

    Teaches idempotency, versioning, and backward compatibility in protocols. Students design protocols that tolerate partial upgrades and retries.

Chapter 3See details

Time, Ordering, and Causality

  • Lesson 1 • Hybrid Logical Clocks

    Combines physical and logical time to achieve causality with bounded skew. Students select the appropriate clock scheme for latency-sensitive applications.

  • Lesson 2 • Physical Clocks and Their Limits

    Examines clock drift, NTP, and GPS-based synchronisation. Motivates logical clocks by showing physical clocks cannot provide global order.

  • Lesson 3 • Consistent Snapshots

    Presents the Chandy-Lamport algorithm for capturing global state. Students apply snapshots to debugging, checkpointing, and deadlock detection.

  • Lesson 4 • Logical Clocks

    Introduces Lamport timestamps and their happens-before relation. Students use logical clocks to order events without physical time.

  • Lesson 5 • Vector Clocks and Version Vectors

    Extends logical clocks to capture causal dependencies per process. Students detect conflicts and reconstruct causal histories using vector clocks.

Chapter 4See details

Replication and Consistency Models

  • Lesson 1 • Conflict Detection and Resolution

    Addresses divergent replicas and strategies for merging concurrent writes. Students implement last-write-wins, CRDTs, and application-level resolution.

  • Lesson 2 • Consistency Models Spectrum

    Defines linearisability, sequential consistency, causal consistency, and eventual consistency. Students map application needs to the appropriate model.

  • Lesson 3 • Quorum-Based Replication

    Explains read and write quorums and their consistency implications. Students calculate quorum sizes for desired durability and availability.

  • Lesson 4 • Replication Goals and Strategies

    Surveys why replication is used and the primary approaches. Frames the consistency-performance trade-off central to this chapter.

  • Lesson 5 • Replication Lag and Read Anomalies

    Identifies stale reads, monotonic read violations, and read-your-writes failures. Students apply session guarantees to eliminate user-visible anomalies.

Chapter 5See details

Consensus and Coordination Protocols

  • Lesson 1 • Distributed Locking and Coordination

    Applies consensus to implement distributed locks, barriers, and leader election services. Students avoid common pitfalls like lock expiry races.

  • Lesson 2 • Byzantine Fault-Tolerant Consensus

    Extends consensus to tolerate malicious or arbitrary failures. Students evaluate when BFT protocols are necessary and their cost.

  • Lesson 3 • Paxos Algorithm

    Walks through single-decree and multi-Paxos phases in detail. Students trace message flows and identify failure recovery paths.

  • Lesson 4 • Raft Consensus Algorithm

    Presents Raft as an understandable alternative to Paxos. Students implement log replication, leader election, and membership changes.

  • Lesson 5 • The Consensus Problem

    Defines consensus formally and explains why it is hard under failures. Connects consensus to replication, locking, and atomic broadcast.

Chapter 6See details

Distributed Storage and Databases

  • Lesson 1 • Partitioning and Sharding

    Covers range, hash, and directory-based partitioning strategies. Students design partition schemes that balance load and minimise cross-shard queries.

  • Lesson 2 • Storage Engine Internals

    Examines LSM trees, B-trees, and write-ahead logs. Students predict performance characteristics and tune storage engines accordingly.

  • Lesson 3 • Distributed Transactions

    Explains two-phase commit, three-phase commit, and their failure modes. Students implement atomic cross-partition operations with correct rollback.

  • Lesson 4 • NoSQL Storage Models

    Surveys key-value, document, column-family, and graph storage models. Students match data models to query patterns and consistency requirements.

  • Lesson 5 • Isolation Levels and Anomalies

    Defines read phenomena and maps them to isolation levels. Students configure isolation to prevent specific anomalies without sacrificing performance.

Chapter 7See details

Fault Tolerance and Resilience Patterns

  • Lesson 1 • Bulkhead and Isolation Patterns

    Isolates failures using thread pools, process boundaries, and resource quotas. Students prevent cascading failures by containing blast radius.

  • Lesson 2 • Chaos Engineering and Resilience Testing

    Introduces controlled fault injection to validate resilience assumptions. Students design experiments that expose hidden failure modes before production.

  • Lesson 3 • Failure Detection

    Covers heartbeats, phi-accrual detectors, and gossip-based detection. Students tune detectors to balance false positives against detection latency.

  • Lesson 4 • Redundancy and Replication Strategies

    Applies active-active, active-passive, and N+1 redundancy models. Students calculate required redundancy for target availability levels.

  • Lesson 5 • Circuit Breakers and Timeouts

    Implements circuit breakers to stop calling failing dependencies. Students configure timeout budgets and half-open state transitions.

Chapter 8See details

Scalability, Performance, and Operations

  • Lesson 1 • Incident Response and Post-mortems

    Establishes on-call practices, runbooks, and blameless post-mortems. Students produce actionable post-mortems that prevent recurrence.

  • Lesson 2 • Capacity Planning and Load Testing

    Forecasts resource needs and validates headroom through load tests. Students build capacity models and interpret load test results.

  • Lesson 3 • Horizontal and Vertical Scaling

    Contrasts scaling strategies and identifies bottlenecks that limit each. Students design stateless services and data tiers that scale independently.

  • Lesson 4 • Deployment and Rolling Updates

    Covers blue-green, canary, and rolling deployment strategies. Students execute zero-downtime deployments and roll back safely on failure.

  • Lesson 5 • Observability: Metrics, Logs, and Traces

    Builds the three pillars of observability into distributed services. Students instrument code and correlate signals to diagnose production issues.

Certification

Your valid completion certificate

This course is for you:

  • Mid-level software engineers: ready to move beyond single-server thinking.

  • Backend developers: hitting the limits of monolithic application architectures.

  • Site reliability engineers: wanting deeper theory behind the systems they operate.

  • Computer science graduates: bridging the gap between coursework and production reality.

  • Platform engineers: designing infrastructure that must survive real-world failure conditions.

  • Tech leads: needing rigorous vocabulary to guide their team's architectural decisions.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast and simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top upskilling courses

FAQ

Who is Dedika?

Is the certificate valid in Australia?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course