Choose your language
Prometheus and Grafana Course
More than 2 million students worldwide

Prometheus and Grafana Course

Master Prometheus and Grafana from installation to production-scale operations. This course takes you through metrics collection, PromQL querying, alerting pipelines, and Grafana visualisation with hands-on, real-world configurations. You will also cover SLOs, distributed tracing, log aggregation with Loki, and Infrastructure as Code for your monitoring stack.

Dedika for businesses

What you'll learn:

You will learn how to install and configure Prometheus, set up exporters for hosts, databases, and applications, and write PromQL queries to extract meaningful metrics. You will build Grafana dashboards with variables, annotations, and drill-down links, and design a complete Alertmanager pipeline with noise-reduction strategies. The course covers service discovery for Kubernetes and cloud computing environments, remote storage integrations, and high-availability deployments. You will also implement SLO-based error budget tracking, add distributed tracing with Jaeger and Tempo, and aggregate logs using Grafana Loki. By the end, you will operate a full observability technology stack confidently in production.

How you study in practice Prometheus and Grafana Course

How you practise Prometheus and Grafana Course

For businesses looking to train their team

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Foundations of Service Monitoring

  • Lesson 1 • Introduction to Prometheus and Grafana

    Positions Prometheus as a metrics engine and Grafana as a visualisation layer. Sets expectations for the full tool stack covered in the course.

  • Lesson 2 • Why Service Monitoring Matters

    Establishes the business and technical case for monitoring. Connects reliability goals to concrete monitoring requirements.

  • Lesson 3 • Observability Pillars Explained

    Defines metrics, logs, and traces as complementary signals. Clarifies how each pillar addresses different failure scenarios.

  • Lesson 4 • Monitoring Architecture Patterns

    Surveys pull-based and push-based collection models. Prepares students to evaluate trade-offs before deploying Prometheus.

Chapter 2See details

Installing and Configuring Prometheus

  • Lesson 1 • Installation Methods and Prerequisites

    Covers binary, Docker, and package-manager installation paths. Students select the method appropriate for their environment.

  • Lesson 2 • Verifying a Running Prometheus Instance

    Uses the built-in UI and API to confirm targets are up and data is flowing. Establishes a baseline health check habit.

  • Lesson 3 • Core Configuration File Structure

    Explains prometheus.yml syntax: global, scrape_configs, and rule_files blocks. Correct configuration is the prerequisite for all subsequent labs.

  • Lesson 4 • Prometheus Architecture Deep Dive

    Maps internal components: retrieval, TSDB, HTTP API, and alerting. Understanding internals prevents misconfiguration in later steps.

  • Lesson 5 • Storage and Retention Settings

    Configures local TSDB retention, chunk encoding, and WAL settings. Students balance disk usage against query performance.

Chapter 3See details

Collecting Metrics with Exporters

  • Lesson 1 • Blackbox Exporter for Synthetic Probing

    Uses Blackbox Exporter to probe HTTP, TCP, DNS, and ICMP endpoints. Provides external availability perspective independent of internal metrics.

  • Lesson 2 • Exporter Concepts and Taxonomy

    Defines what an exporter is and how it exposes the /metrics endpoint. Classifies exporters by target type to guide selection.

  • Lesson 3 • Application-Level Instrumentation

    Integrates Prometheus client libraries into application code. Students expose custom business and performance metrics.

  • Lesson 4 • Database and Middleware Exporters

    Configures exporters for relational databases, caches, and message brokers. Extends visibility beyond host-level data.

  • Lesson 5 • Node Exporter for Host Metrics

    Deploys Node Exporter to collect CPU, memory, disk, and network data. These host metrics form the baseline for infrastructure dashboards.

Chapter 4See details

Querying Metrics with PromQL

  • Lesson 1 • PromQL Data Types and Selectors

    Introduces instant vectors, range vectors, scalars, and strings. Correct type usage prevents query errors in all downstream work.

  • Lesson 2 • Rate, Increase, and Delta Functions

    Converts counters into per-second rates and computes changes over time. These functions are the foundation of throughput and error-rate queries.

  • Lesson 3 • Histogram and Summary Metrics

    Queries latency distributions using histogram_quantile and summary quantiles. Enables SLO-aligned percentile calculations.

  • Lesson 4 • Advanced PromQL Techniques

    Combines binary operators, vector matching, and subqueries for complex analysis. Prepares students for production-grade alerting expressions.

  • Lesson 5 • Aggregation and Grouping Operators

    Applies sum, avg, max, min, count, and topk across label dimensions. Aggregation is essential for fleet-wide and per-service views.

Chapter 5See details

Alerting with Prometheus and Alertmanager

  • Lesson 1 • Routing, Grouping, and Inhibition

    Configures route trees to direct alerts to the right receivers. Grouping and inhibition reduce noise during large-scale incidents.

  • Lesson 2 • Alert Quality and Runbook Design

    Applies SLO-based thresholds and multi-window burn-rate alerts. Runbook annotations guide on-call engineers to fast resolution.

  • Lesson 3 • Alerting Rule Fundamentals

    Defines alert rules using PromQL expressions, for durations, and labels. Rules are the trigger mechanism for the entire alerting pipeline.

  • Lesson 4 • Notification Receivers and Templates

    Integrates email, Slack, PagerDuty, and webhook receivers. Custom templates make notifications informative and actionable.

  • Lesson 5 • Installing and Configuring Alertmanager

    Deploys Alertmanager and connects it to Prometheus via the alerting block. Correct wiring ensures alerts reach the notification layer.

Chapter 6See details

Visualising Metrics with Grafana

  • Lesson 1 • Dashboard and Panel Fundamentals

    Creates dashboards with time-series, stat, gauge, and table panels. Panel types are matched to the nature of each metric.

  • Lesson 2 • Variables and Template Queries

    Adds dashboard variables backed by label_values queries for dynamic filtering. Variables transform static dashboards into reusable fleet-wide tools.

  • Lesson 3 • Dashboard Best Practices and Governance

    Applies naming conventions, layout principles, and version control for dashboards. Governance prevents dashboard sprawl in team environments.

  • Lesson 4 • Annotations, Links, and Drill-Downs

    Overlays deployment events as annotations and links panels to related dashboards. Context enrichment accelerates root-cause analysis.

  • Lesson 5 • Grafana Installation and Initial Setup

    Installs Grafana and connects it to a Prometheus data source. A working data source connection is required for all dashboard work.

Chapter 7See details

Service Discovery and Label Management

  • Lesson 1 • Relabelling for Label Normalisation

    Uses relabel_configs to rename, drop, and transform labels before ingestion. Consistent labels are required for accurate aggregation and alerting.

  • Lesson 2 • File-Based and DNS Service Discovery

    Implements file_sd_configs and dns_sd_configs for lightweight automation. These methods work without cloud provider APIs.

  • Lesson 3 • Metric Relabelling and Cardinality Control

    Applies metric_relabel_configs to drop high-cardinality series at scrape time. Cardinality control protects TSDB performance and storage.

  • Lesson 4 • Static vs. Dynamic Service Discovery

    Contrasts static_configs with automated discovery mechanisms. Dynamic discovery is essential for containerised and cloud-native workloads.

  • Lesson 5 • Cloud and Container Service Discovery

    Configures Kubernetes, EC2, and Consul service discovery. Covers the most common production environments for Prometheus deployments.

Chapter 8See details

Production Operations and Scalability

  • Lesson 1 • Performance Tuning and Capacity Planning

    Profiles Prometheus resource usage and forecasts storage growth. Capacity planning prevents outages caused by monitoring infrastructure itself.

  • Lesson 2 • Federation and Hierarchical Scraping

    Configures Prometheus federation to aggregate metrics across clusters. Federation enables global views without centralising all raw data.

  • Lesson 3 • Remote Storage Integrations

    Connects Prometheus to long-term storage backends via remote_write and remote_read. Extends retention beyond local TSDB limits.

  • Lesson 4 • Security Hardening and Access Control

    Enables TLS, basic auth, and Grafana role-based access control. Security hardening is mandatory before exposing monitoring to production networks.

  • Lesson 5 • High Availability for Prometheus

    Runs duplicate Prometheus instances with deduplication at the query layer. HA prevents monitoring blind spots during instance failures.

Certification

Your valid completion certificate

This course is for you:

  • Backend Developer: wants to stop flying blind when services degrade.

  • DevOps Engineer: needs to replace ad-hoc monitoring with a structured observability stack.

  • Site Reliability Engineer: ready to formalise SLO tracking and on-call workflows.

  • Systems Administrator: transitioning from legacy monitoring tools to cloud-native solutions.

  • Platform Engineer: responsible for giving development teams reliable, self-service dashboards.

  • Software Engineering Student: building job-ready skills in production observability tooling.

What our students say

Your lessons are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to change platforms... I'm grateful for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can change chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the way videos are presented and transcribed, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast and simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top upskilling courses

FAQ

Who is Dedika?

Is the certificate valid in Australia?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course