Choose your language
Site Reliability Engineering Course
More than 2 million students worldwide

Site Reliability Engineering Course

4.8

Master the engineering discipline that keeps the world's most critical systems running. This course gives you the frameworks, tools, and hands-on practices used by SRE teams at scale — from SLOs and error budgets to incident command and chaos engineering. Whether you're moving into SRE or levelling up your current role, this is the structured foundation you need.

Dedika for businesses

What you'll learn:

You'll build a complete SRE skill set starting with core principles, SLIs, SLOs, and error budgets, then move into observability, alerting design, and on-call management. You'll learn how to lead incident response, run blameless postmortems, and drive continuous improvement. The course covers toil reduction, infrastructure automation, progressive delivery, and capacity planning. You'll also explore resilience engineering, chaos experiments, Kubernetes reliability, and AIOps techniques. By the end, you'll have the technical depth and organisational skills to build and scale reliable systems.

How you study in practice Site Reliability Engineering Course

How you practise Site Reliability Engineering Course

For businesses looking to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 38 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Site Reliability Engineering

  • Lesson 1 • SRE Team Structure and Roles

    Examines how SRE teams are organised, staffed, and embedded within organisations. Clarifies role boundaries with development and platform teams.

  • Lesson 2 • Key SRE Concepts and Terminology

    Defines SLIs, SLOs, SLAs, error budgets, and toil. Shared vocabulary enables precise communication across engineering and business teams.

  • Lesson 3 • Reliability Culture and Incentives

    Explores psychological safety, blameless culture, and incentive alignment. Cultural foundations determine whether technical practices succeed or fail.

  • Lesson 4 • Origins and Philosophy of SRE

    Traces SRE's emergence from software engineering applied to operations. Provides historical context that frames every subsequent technical practice.

Chapter 2See details

Measuring Reliability with SLIs and SLOs

  • Lesson 1 • Error Budget Policy and Management

    Defines how error budgets are calculated, tracked, and enforced. Error budgets translate abstract reliability goals into actionable engineering decisions.

  • Lesson 2 • Setting Realistic SLO Targets

    Covers negotiation, baselining, and iterative refinement of SLO values. Targets must balance user expectations with engineering feasibility.

  • Lesson 3 • Choosing Meaningful SLIs

    Guides selection of metrics that accurately reflect user experience. Poor SLI choice undermines the entire reliability measurement system.

  • Lesson 4 • SLO Dashboards and Reporting

    Teaches construction of SLO dashboards and automated reporting pipelines. Visibility ensures teams act on reliability data rather than ignore it.

Chapter 3See details

Monitoring, Observability, and Alerting

  • Lesson 1 • On-Call Management and Runbooks

    Structures on-call rotations, escalation paths, and runbook creation. Effective on-call practices reduce mean time to resolution and engineer burnout.

  • Lesson 2 • Observability Pillars: Metrics, Logs, Traces

    Introduces the three observability pillars and their complementary roles. Understanding each pillar prevents over-reliance on any single data source.

  • Lesson 3 • Monitoring System Architecture

    Examines scalable monitoring architectures including push vs. pull models and long-term storage. Architecture choices affect cost, reliability, and query performance.

  • Lesson 4 • Instrumentation and Data Collection

    Covers code-level and infrastructure-level instrumentation techniques. Proper instrumentation is the prerequisite for meaningful observability.

  • Lesson 5 • Effective Alerting Design

    Teaches symptom-based alerting, multi-window burn rate alerts, and alert routing. Well-designed alerts reduce on-call fatigue and improve response time.

Chapter 4See details

Incident Management and Response

  • Lesson 1 • Diagnosis and Mitigation Techniques

    Teaches systematic diagnosis using observability data and hypothesis-driven debugging. Fast mitigation limits user impact whilst root cause analysis continues.

  • Lesson 2 • Incident Command and Coordination

    Covers incident commander, communications lead, and scribe roles. Clear role assignment prevents confusion and speeds resolution during high-stress events.

  • Lesson 3 • Blameless Postmortems

    Guides writing and facilitating blameless postmortems that produce actionable improvements. Postmortems are the primary mechanism for organisational learning.

  • Lesson 4 • Incident Classification and Severity

    Defines severity levels, impact criteria, and triage procedures. Consistent classification ensures proportional response and resource allocation.

  • Lesson 5 • Incident Metrics and Continuous Improvement

    Tracks MTTD, MTTR, and incident frequency to measure response effectiveness. Trend analysis drives targeted improvements to processes and tooling.

Chapter 5See details

Toil Reduction and Automation

  • Lesson 1 • Automation Design Principles

    Covers idempotency, safety, and observability requirements for operational automation. Poorly designed automation creates new failure modes worse than the toil it replaces.

  • Lesson 2 • Measuring Automation Effectiveness

    Tracks toil reduction, automation coverage, and failure rates of automated systems. Measurement validates automation investment and surfaces regressions.

  • Lesson 3 • Automating Operational Workflows

    Applies automation to common SRE tasks such as deployments, scaling, and remediation. Workflow automation frees engineers for higher-value reliability work.

  • Lesson 4 • Infrastructure as Code Fundamentals

    Introduces declarative infrastructure provisioning and configuration management. IaC is the foundation for reproducible, auditable infrastructure automation.

  • Lesson 5 • Identifying and Quantifying Toil

    Defines toil characteristics and measurement methods for tracking its cost. Quantification builds the business case for automation investment.

Chapter 6See details

Release Engineering and Change Management

  • Lesson 1 • Feature Flags and Dark Launches

    Covers feature flag architecture, targeting rules, and operational hygiene. Flags decouple deployment from release, enabling safer experimentation.

  • Lesson 2 • Release Engineering Principles

    Establishes reproducibility, hermetic builds, and policy-as-code for releases. Sound release engineering prevents configuration drift and supply chain issues.

  • Lesson 3 • Continuous Integration and Delivery

    Covers CI pipeline design, test gates, and deployment frequency optimisation. High deployment frequency with quality gates reduces batch size and blast radius.

  • Lesson 4 • Rollback and Rollforward Strategies

    Defines criteria and mechanisms for fast rollback and rollforward decisions. Reliable rollback capability is the safety net for all deployment strategies.

  • Lesson 5 • Progressive Delivery Strategies

    Teaches canary releases, blue-green deployments, and traffic splitting. Progressive delivery limits blast radius and enables data-driven rollout decisions.

Chapter 7See details

Capacity Planning and Performance Engineering

  • Lesson 1 • Capacity Modelling Techniques

    Introduces resource utilisation models, queuing theory basics, and headroom targets. Models translate demand forecasts into concrete provisioning decisions.

  • Lesson 2 • Load Testing and Benchmarking

    Teaches load test design, execution, and result interpretation for capacity validation. Load tests reveal bottlenecks before production traffic exposes them.

  • Lesson 3 • Demand Forecasting and Load Modelling

    Covers traffic pattern analysis, seasonality modelling, and growth projection. Accurate demand forecasts prevent both under-provisioning and over-spending.

  • Lesson 4 • Auto-Scaling and Elastic Infrastructure

    Designs horizontal and vertical auto-scaling policies tied to SLO-relevant signals. Elastic infrastructure absorbs demand spikes whilst controlling cost.

  • Lesson 5 • Performance Optimisation Strategies

    Covers profiling, caching, connection pooling, and query optimisation. Performance improvements extend capacity without additional infrastructure spend.

Chapter 8See details

Resilience Engineering and Fault Tolerance

  • Lesson 1 • Chaos Engineering Fundamentals

    Introduces chaos engineering principles, hypothesis design, and blast radius control. Controlled failure injection reveals hidden weaknesses before real incidents do.

  • Lesson 2 • Redundancy and Failover Architecture

    Examines active-active, active-passive, and multi-region redundancy models. Redundancy architecture determines recovery time and data loss objectives.

  • Lesson 3 • Resilience Design Patterns

    Covers circuit breakers, bulkheads, retries, and timeouts as core resilience primitives. Patterns prevent cascading failures and limit blast radius.

  • Lesson 4 • Dependency and Risk Management

    Maps external dependencies, assesses failure probability, and designs mitigation controls. Dependency risk is a leading cause of reliability incidents.

  • Lesson 5 • Disaster Recovery Planning

    Covers DR strategy selection, runbook creation, and regular DR testing cadence. Untested DR plans fail when needed most.

Certification

Your valid completion certificate

This course is for you:

  • DevOps engineers: ready to adopt a more rigorous, reliability-focused engineering discipline.

  • Backend developers: taking on operational responsibilities and needing a structured approach.

  • Systems administrators: transitioning from traditional ops into modern SRE practices.

  • Platform engineers: looking to formalise reliability standards across multiple development teams.

  • Engineering managers: building or overseeing an SRE function for the first time.

  • Career changers: coming from IT or QA backgrounds and targeting SRE roles specifically.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast and simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top upskilling courses

FAQ

Who is Dedika?

Is the certificate valid in Australia?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course