AzureINTERMEDIATE

Observability & Monitoring Engineering

Master production-grade cloud and DevOps skills with hands-on labs and professional roadmaps.

1learners

What you'll learn

Implement production-grade Cloud architecture
Master Site Reliability Engineering (SRE) patterns
Build automated CI/CD pipelines for Kubernetes
Implement centralized observability & monitoring
Apply security-first (DevSecOps) principles
Optimize cloud costs and platform reliability

Course Description

The Observability & Monitoring Engineering pathway is for engineers who want to be the person every team in the organisation turns to when nobody can figure out what is wrong. Observability engineers are the architects of visibility — the professionals who build the systems that tell the business exactly what its software is doing, when it breaks, why it broke, and how to fix it. In 2025, as distributed systems become more complex and the cost of downtime rises, this is one of the fastest-growing, highest-impact specialisations in cloud engineering.

This pathway is built around a simple but powerful premise: monitoring tells you something is wrong. Observability tells you why. Most organisations have monitoring — dashboards, alerts, and on-call rotations. Very few have genuine observability — the ability to ask arbitrary questions about their systems and get answers, even questions they did not anticipate needing to ask. This programme teaches participants to build both, at enterprise scale, across the most in-demand toolchains in the industry.

Learning Outcomes

By the end of this programme, participants will be able to:

  • Design and deploy a production-grade open-source LGTM observability stack (Loki, Grafana, Tempo, Mimir) on Kubernetes

  • Instrument applications in any language using OpenTelemetry SDKs and configure the OpenTelemetry Collector as an enterprise telemetry pipeline

  • Build production-grade Prometheus monitoring with PromQL expertise, recording rules, AlertManager routing, and Thanos for long-term scalability

  • Master Grafana deeply — advanced dashboard engineering, templating, alerting, Grafana OnCall, and dashboards-as-code with Grafonnet and Jsonnet

  • Implement and operate Datadog as a complete enterprise observability platform — APM, logs, metrics, RUM, Synthetics, and monitors-as-code with Terraform

  • Configure and operate Dynatrace for AI-powered full-stack observability and anomaly detection

  • Build and manage the ELK Stack (Elasticsearch, Logstash, Kibana) for enterprise log management at scale

  • Instrument distributed systems with distributed tracing — understanding trace context propagation, sampling strategies, and tail-based sampling

  • Build comprehensive AWS CloudWatch observability — custom metrics, Container Insights, X-Ray, Synthetics, and cross-account monitoring

  • Implement AIOps using ML-based anomaly detection, intelligent alerting, and AI-assisted root cause analysis

  • Design observability cost governance strategies — cardinality reduction, log filtering, sampling, and telemetry pipeline optimisation

  • Build everything as code — dashboards, alerts, SLOs, and monitoring configurations defined in Git and deployed via CI/CD

Course Curriculum

9 Sections · 0 Lessons
1

OpenTelemetry — The Universal Instrumentation Standard

0 Lessons
2

Prometheus — Metrics at Scale

0 Lessons
3

Grafana — Dashboard Engineering & Unified Observability

0 Lessons
4

Log Management — Loki, ELK Stack & Enterprise Log Analytics

0 Lessons
5

Distributed Tracing — End-to-End Request Visibility

0 Lessons
6

Commercial Observability Platforms — Datadog, Dynatrace & New Relic

0 Lessons
7

AWS Native Observability — CloudWatch, X-Ray & Container Insights

0 Lessons
8

AIOps, Continuous Profiling & Emerging Observability

0 Lessons
9

Observability at Scale — Cost, Governance & Platform Engineering

0 Lessons

Ratings & Reviews

0 reviews

0.0
Leave a review
You must be enrolled to submit.
Submitting updates your previous review.