
Prometheus and Grafana Course
Master Prometheus and Grafana from installation to production-scale operations. This course takes you through metrics collection, PromQL querying, alerting pipelines, and Grafana visualisation with hands-on, real-world configurations. You'll also cover SLOs, distributed tracing, log aggregation with Loki, and Infrastructure as Code for your monitoring stack.
What you will learn:
You will learn how to install and configure Prometheus, set up exporters for hosts, databases, and applications, and write PromQL queries to extract meaningful metrics. You will build Grafana dashboards with variables, annotations, and drill-down links, and design a complete Alertmanager pipeline with noise-reduction strategies. The course covers service discovery for Kubernetes and cloud environments, remote storage integrations, and high-availability deployments. You will also implement SLO-based error budget tracking, add distributed tracing with Jaeger and Tempo, and aggregate logs using Grafana Loki. By the end, you will operate a full observability stack confidently in production.
How you study in practice Prometheus and Grafana Course
How you practise Prometheus and Grafana Course
For companies looking to train their teams
With Dedika for businesses, the course includes exercises and examples tailored to your company and its specific needs.
Course content
8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)
Chapter 1HideHide detailsSee detailsFoundations of Service Monitoring
Foundations of Service Monitoring
Lesson 1 • Introduction to Prometheus and Grafana
Positions Prometheus as a metrics engine and Grafana as a visualisation layer. Sets expectations for the full tool stack covered in the course.
Lesson 2 • Why Service Monitoring Matters
Establishes the business and technical case for monitoring. Connects reliability goals to concrete monitoring requirements.
Lesson 3 • Observability Pillars Explained
Defines metrics, logs, and traces as complementary signals. Clarifies how each pillar addresses different failure scenarios.
Lesson 4 • Monitoring Architecture Patterns
Surveys pull-based and push-based collection models. Prepares students to evaluate trade-offs before deploying Prometheus.
Chapter 2HideHide detailsSee detailsInstalling and Configuring Prometheus
Installing and Configuring Prometheus
Lesson 1 • Installation Methods and Prerequisites
Covers binary, Docker, and package-manager installation paths. Students select the method appropriate for their environment.
Lesson 2 • Verifying a Running Prometheus Instance
Uses the built-in UI and API to confirm targets are up and data is flowing. Establishes a baseline health check habit.
Lesson 3 • Core Configuration File Structure
Explains prometheus.yml syntax: global, scrape_configs, and rule_files blocks. Correct configuration is the prerequisite for all subsequent labs.
Lesson 4 • Prometheus Architecture Deep Dive
Maps internal components: retrieval, TSDB, HTTP API, and alerting. Understanding internals prevents misconfiguration in later steps.
Lesson 5 • Storage and Retention Settings
Configures local TSDB retention, chunk encoding, and WAL settings. Students balance disk usage against query performance.
Chapter 3HideHide detailsSee detailsCollecting Metrics with Exporters
Collecting Metrics with Exporters
Lesson 1 • Blackbox Exporter for Synthetic Probing
Uses Blackbox Exporter to probe HTTP, TCP, DNS, and ICMP endpoints. Provides external availability perspective independent of internal metrics.
Lesson 2 • Exporter Concepts and Taxonomy
Defines what an exporter is and how it exposes the /metrics endpoint. Classifies exporters by target type to guide selection.
Lesson 3 • Application-Level Instrumentation
Integrates Prometheus client libraries into application code. Students expose custom business and performance metrics.
Lesson 4 • Database and Middleware Exporters
Configures exporters for relational databases, caches, and message brokers. Extends visibility beyond host-level data.
Lesson 5 • Node Exporter for Host Metrics
Deploys Node Exporter to collect CPU, memory, disk, and network data. These host metrics form the baseline for infrastructure dashboards.
Chapter 4HideHide detailsSee detailsQuerying Metrics with PromQL
Querying Metrics with PromQL
Lesson 1 • PromQL Data Types and Selectors
Introduces instant vectors, range vectors, scalars, and strings. Correct type usage prevents query errors in all downstream work.
Lesson 2 • Rate, Increase, and Delta Functions
Converts counters into per-second rates and computes changes over time. These functions are the foundation of throughput and error-rate queries.
Lesson 3 • Histogram and Summary Metrics
Queries latency distributions using histogram_quantile and summary quantiles. Enables SLO-aligned percentile calculations.
Lesson 4 • Advanced PromQL Techniques
Combines binary operators, vector matching, and subqueries for complex analysis. Prepares students for production-grade alerting expressions.
Lesson 5 • Aggregation and Grouping Operators
Applies sum, avg, max, min, count, and topk across label dimensions. Aggregation is essential for fleet-wide and per-service views.
Chapter 5HideHide detailsSee detailsAlerting with Prometheus and Alertmanager
Alerting with Prometheus and Alertmanager
Lesson 1 • Routing, Grouping, and Inhibition
Configures route trees to direct alerts to the right receivers. Grouping and inhibition reduce noise during large-scale incidents.
Lesson 2 • Alert Quality and Runbook Design
Applies SLO-based thresholds and multi-window burn-rate alerts. Runbook annotations guide on-call engineers to fast resolution.
Lesson 3 • Alerting Rule Fundamentals
Defines alert rules using PromQL expressions, for durations, and labels. Rules are the trigger mechanism for the entire alerting pipeline.
Lesson 4 • Notification Receivers and Templates
Integrates email, Slack, PagerDuty, and webhook receivers. Custom templates make notifications informative and actionable.
Lesson 5 • Installing and Configuring Alertmanager
Deploys Alertmanager and connects it to Prometheus via the alerting block. Correct wiring ensures alerts reach the notification layer.
Chapter 6HideHide detailsSee detailsVisualising Metrics with Grafana
Visualising Metrics with Grafana
Lesson 1 • Dashboard and Panel Fundamentals
Creates dashboards with time-series, stat, gauge, and table panels. Panel types are matched to the nature of each metric.
Lesson 2 • Variables and Template Queries
Adds dashboard variables backed by label_values queries for dynamic filtering. Variables transform static dashboards into reusable fleet-wide tools.
Lesson 3 • Dashboard Best Practices and Governance
Applies naming conventions, layout principles, and version control for dashboards. Governance prevents dashboard sprawl in team environments.
Lesson 4 • Annotations, Links, and Drill-Downs
Overlays deployment events as annotations and links panels to related dashboards. Context enrichment accelerates root-cause analysis.
Lesson 5 • Grafana Installation and Initial Setup
Installs Grafana and connects it to a Prometheus data source. A working data source connection is required for all dashboard work.
Chapter 7HideHide detailsSee detailsService Discovery and Label Management
Service Discovery and Label Management
Lesson 1 • Relabelling for Label Normalisation
Uses relabel_configs to rename, drop, and transform labels before ingestion. Consistent labels are required for accurate aggregation and alerting.
Lesson 2 • File-Based and DNS Service Discovery
Implements file_sd_configs and dns_sd_configs for lightweight automation. These methods work without cloud provider APIs.
Lesson 3 • Metric Relabelling and Cardinality Control
Applies metric_relabel_configs to drop high-cardinality series at scrape time. Cardinality control protects TSDB performance and storage.
Lesson 4 • Static vs. Dynamic Service Discovery
Contrasts static_configs with automated discovery mechanisms. Dynamic discovery is essential for containerised and cloud-native workloads.
Lesson 5 • Cloud and Container Service Discovery
Configures Kubernetes, EC2, and Consul service discovery. Covers the most common production environments for Prometheus deployments.
Chapter 8HideHide detailsSee detailsProduction Operations and Scalability
Production Operations and Scalability
Lesson 1 • Performance Tuning and Capacity Planning
Profiles Prometheus resource usage and forecasts storage growth. Capacity planning prevents outages caused by monitoring infrastructure itself.
Lesson 2 • Federation and Hierarchical Scraping
Configures Prometheus federation to aggregate metrics across clusters. Federation enables global views without centralising all raw data.
Lesson 3 • Remote Storage Integrations
Connects Prometheus to long-term storage backends via remote_write and remote_read. Extends retention beyond local TSDB limits.
Lesson 4 • Security Hardening and Access Control
Enables TLS, basic auth, and Grafana role-based access control. Security hardening is mandatory before exposing monitoring to production networks.
Lesson 5 • High Availability for Prometheus
Runs duplicate Prometheus instances with deduplication at the query layer. HA prevents monitoring blind spots during instance failures.
Your valid completion certificate
This course is for you:
Backend Developer: wants to stop flying blind when services degrade.
DevOps Engineer: needs to replace ad-hoc monitoring with a structured observability stack.
Site Reliability Engineer: ready to formalise SLO tracking and on-call workflows.
Systems Administrator: transitioning from legacy monitoring tools to cloud-native solutions.
Platform Engineer: responsible for giving development teams reliable, self-service dashboards.
Software Engineering Student: building job-ready skills in production observability tooling.
What our students say
Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...

I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.

I like the content and the way videos are presented and transcribed, which speeds up the process!

The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.

Top qualifications
FAQ
Who is Dedika?
Is the certificate valid in South Africa?
Are the courses free?
What is the course workload?
What are the courses like?
How do the courses work?
What is the duration of the courses?
What is the cost or price of the courses?
What is an EAD or online course and how does it work?
PDF Course




















