Choose your language
Prometheus Monitoring Training
More than 2 million students worldwide

Prometheus Monitoring Training

Master Prometheus from installation to production-scale deployment in this comprehensive monitoring training. You'll configure scraping, write PromQL queries, build Grafana dashboards, and design reliable alerting pipelines. By the end, you'll have the hands-on skills to monitor any infrastructure with confidence.

Dedika for businesses

What you will learn:

This course covers every layer of the Prometheus monitoring stack, starting with core concepts like metric types and the pull model, then moving into installation, service discovery, and exporters. You will develop strong PromQL skills to query and aggregate time-series data for dashboards and alerts. You will configure Alertmanager routing, grouping, and notification channels to build a low-noise alerting pipeline. The course also covers Grafana dashboard design, long-term storage backends like Thanos and Mimir, Kubernetes monitoring with the Prometheus Operator, and security hardening. You will finish with practical knowledge of SLO management, capacity planning, and incident response workflows.

How you study in practice Prometheus Monitoring Training

How you practice Prometheus Monitoring Training

For companies looking to train their teams

With Dedika for businesses, the course includes exercises and examples tailored to your own business and the way your company needs.

Click here

Course Content

8 Chapters • 39 LessonsDuration between 4 and 360 hours (you decide)

Chapter 1See details

Introduction to Prometheus and Monitoring

  • Lesson 1 • Prometheus History and Design Goals

    Traces Prometheus's origin at SoundCloud and its CNCF adoption. Explains design decisions that shape every configuration choice students will make.

  • Lesson 2 • Prometheus Architecture Overview

    Maps all major components: server, exporters, Alertmanager, and Pushgateway. Students gain a mental model that guides installation and configuration decisions.

  • Lesson 3 • Observability and Monitoring Fundamentals

    Defines monitoring, observability, and the three pillars: metrics, logs, and traces. Establishes the conceptual baseline for all Prometheus topics ahead.

  • Lesson 4 • Metric Types and Data Model

    Covers Counter, Gauge, Histogram, and Summary types with their appropriate use cases. Correct type selection is foundational to accurate querying later.

Chapter 2See details

Installing and Configuring Prometheus

  • Lesson 1 • Storage and Retention Settings

    Explains the local TSDB, block compaction, and retention flags. Students tune storage to balance disk usage against query performance.

  • Lesson 2 • Installation Methods and Prerequisites

    Covers binary, Docker, and package-manager installation paths. Students choose the method that fits their environment before proceeding to configuration.

  • Lesson 3 • Running Prometheus as a Service

    Covers systemd unit files, startup flags, and log management. Students leave with a production-grade service that survives reboots.

  • Lesson 4 • Scrape Configuration and Targets

    Teaches how to define scrape jobs, set intervals, and apply per-job overrides. Students configure Prometheus to collect metrics from real endpoints.

  • Lesson 5 • Core Configuration File Structure

    Explains prometheus.yml sections: global, scrape_configs, and rule_files. Mastery here prevents the most common misconfiguration errors.

Chapter 3See details

Service Discovery and Target Management

  • Lesson 1 • File-Based Service Discovery

    Configures file_sd_configs with JSON and YAML target files and hot-reload behavior. Students automate target updates without restarting Prometheus.

  • Lesson 2 • Cloud and Orchestrator Discovery

    Covers Kubernetes, EC2, Consul, and other built-in discovery integrations. Students connect Prometheus to dynamic environments common in production.

  • Lesson 3 • Validating and Troubleshooting Targets

    Uses the /targets UI page, HTTP API, and logs to diagnose scrape failures. Students resolve common discovery and connectivity issues independently.

  • Lesson 4 • Relabeling for Target Transformation

    Explains relabel_configs actions: replace, keep, drop, labelmap, and hashmod. Relabeling is the primary tool for shaping target metadata and filtering.

  • Lesson 5 • Static vs. Dynamic Service Discovery

    Contrasts static file targets with dynamic discovery and explains when each is appropriate. Sets the motivation for learning discovery integrations.

Chapter 4See details

Exporters and Instrumentation

  • Lesson 1 • Official Exporters Overview

    Surveys Node Exporter, Blackbox Exporter, and other widely used official exporters. Students understand which exporter fits each monitoring target.

  • Lesson 2 • Blackbox Exporter for Synthetic Probing

    Configures HTTP, TCP, ICMP, and DNS probes to monitor external endpoints. Students detect availability and latency issues from an outside perspective.

  • Lesson 3 • Deploying and Configuring Node Exporter

    Walks through installing Node Exporter, enabling collectors, and securing the endpoint. Node Exporter is the most common starting point for infrastructure monitoring.

  • Lesson 4 • Custom Instrumentation with Client Libraries

    Teaches metric registration and exposition using Go, Python, and Java client libraries. Students add Prometheus metrics directly to application code.

  • Lesson 5 • Pushgateway for Batch Jobs

    Explains when to use Pushgateway, its limitations, and correct push patterns. Students avoid common misuse that leads to stale or misleading metrics.

Chapter 5See details

PromQL Query Language Fundamentals

  • Lesson 1 • Arithmetic and Comparison Operators

    Covers binary arithmetic, comparison, and logical operators with vector matching. Students combine metrics from different series correctly.

  • Lesson 2 • Advanced PromQL Techniques

    Covers predict_linear, histogram_quantile, label_replace, and subqueries. Students solve complex real-world query problems with these advanced tools.

  • Lesson 3 • Aggregation Operators

    Teaches sum, avg, min, max, count, and topk with by and without clauses. Aggregation is essential for reducing cardinality in dashboards and alerts.

  • Lesson 4 • Functions for Rate and Change

    Explains rate, irate, increase, delta, and deriv for analyzing counter and gauge trends. Students select the correct function for each metric type.

  • Lesson 5 • PromQL Data Types and Selectors

    Introduces instant vectors, range vectors, scalars, and string types. Students understand what each query returns before applying functions.

Chapter 6See details

Alerting with Prometheus and Alertmanager

  • Lesson 1 • Inhibition, Silencing, and Alert Quality

    Uses inhibit_rules to suppress downstream alerts and silences for maintenance windows. Students reduce alert fatigue while preserving signal integrity.

  • Lesson 2 • Receivers and Notification Channels

    Sets up email, Slack, PagerDuty, webhook, and other receiver integrations. Students configure and test each channel to confirm delivery.

  • Lesson 3 • Routing and Grouping Alerts

    Configures route trees, group_by, group_wait, and group_interval settings. Proper routing ensures the right team receives the right alert at the right time.

  • Lesson 4 • Alertmanager Architecture and Setup

    Explains Alertmanager's pipeline: routing, grouping, inhibition, and silencing. Students install and connect Alertmanager to Prometheus correctly.

  • Lesson 5 • Alerting Rules and Recording Rules

    Defines rule file syntax, alert expressions, labels, and annotations. Recording rules pre-compute expensive queries to improve alert evaluation speed.

Chapter 7See details

Visualization with Grafana and Prometheus

  • Lesson 1 • Connecting Grafana to Prometheus

    Adds Prometheus as a Grafana data source and verifies connectivity. This is the prerequisite step before any visualization work can begin.

  • Lesson 2 • Dashboard Design Best Practices

    Applies USE and RED method frameworks to structure dashboards for fast incident response. Students evaluate and improve existing dashboards against these standards.

  • Lesson 3 • Variables and Templating

    Defines query, custom, and interval variables to make dashboards dynamic. Templating allows a single dashboard to cover multiple services or environments.

  • Lesson 4 • Importing and Managing Dashboards

    Imports community dashboards from Grafana.com and manages them with version control. Students avoid dashboard sprawl through folder organization and permissions.

  • Lesson 5 • Building Panels and Dashboards

    Creates time-series, stat, gauge, table, and heatmap panels with PromQL queries. Students assemble panels into coherent dashboards for a specific service.

Chapter 8See details

Scaling, Federation, and Long-Term Storage

  • Lesson 1 • Long-Term Storage Backends

    Surveys Thanos, Cortex, Mimir, and VictoriaMetrics as remote storage solutions. Students select a backend based on retention, query, and operational requirements.

  • Lesson 2 • Scaling Prometheus Horizontally

    Addresses cardinality limits, sharding strategies, and functional decomposition. Students identify when a single Prometheus instance is insufficient.

  • Lesson 3 • Remote Write and Remote Read

    Configures remote_write and remote_read to integrate external storage systems. Students offload long-term retention from local TSDB to scalable backends.

  • Lesson 4 • Prometheus Federation

    Configures hierarchical and cross-service federation using the /federate endpoint. Students aggregate metrics from multiple Prometheus servers into a global view.

  • Lesson 5 • High Availability for Prometheus

    Runs redundant Prometheus pairs and uses Alertmanager clustering to eliminate single points of failure. Students verify that deduplication works correctly end to end.

Certification

Your valid completion certificate

This course is for you:

  • DevOps engineer: wants deeper observability skills beyond basic log monitoring.

  • Site reliability engineer: needs structured alerting and SLO tracking at scale.

  • Backend developer: ready to instrument applications and understand production behavior.

  • Systems administrator: transitioning from legacy monitoring tools to cloud-native stacks.

  • Platform engineer: building internal tooling and needs a reliable metrics foundation.

  • Career changer: moving into infrastructure roles and wants job-ready monitoring expertise.

What our students say

Your classes are perfect. I purchased the one-year package and finally have the opportunity to follow various topics of interest without needing to switch platforms... I thank you for everything you do, I've already recommended you to other people...
Giulio Carlo
Giulio CarloDigital Marketing Student
I like how the lessons are straight to the point and how I can switch chapters and skip content I don't need.
Mariana Ferres
Mariana FerresPhotography Student
I like the content and the presentation style and video transcription, which speeds up the process!
Luciana Alvarenga
Luciana AlvarengaNail Design Student
The platform is fast, simple to use. The diversity of content and complementary videos really help with learning.
André Felipe
André FelipePrompt Engineering Student

Top trainings

FAQ

Who is Dedika?

Is the certificate valid in United States?

Are the courses free?

What is the course workload?

What are the courses like?

How do the courses work?

What is the duration of the courses?

What is the cost or price of the courses?

What is an EAD or online course and how does it work?

PDF Course