Observability & Monitoring Engineering
Master production-grade cloud and DevOps skills with hands-on labs and professional roadmaps.
What you'll learn
Course Description
The Observability & Monitoring Engineering pathway is for engineers who want to be the person every team in the organisation turns to when nobody can figure out what is wrong. Observability engineers are the architects of visibility — the professionals who build the systems that tell the business exactly what its software is doing, when it breaks, why it broke, and how to fix it. In 2025, as distributed systems become more complex and the cost of downtime rises, this is one of the fastest-growing, highest-impact specialisations in cloud engineering.
This pathway is built around a simple but powerful premise: monitoring tells you something is wrong. Observability tells you why. Most organisations have monitoring — dashboards, alerts, and on-call rotations. Very few have genuine observability — the ability to ask arbitrary questions about their systems and get answers, even questions they did not anticipate needing to ask. This programme teaches participants to build both, at enterprise scale, across the most in-demand toolchains in the industry.
Learning Outcomes
By the end of this programme, participants will be able to:
Design and deploy a production-grade open-source LGTM observability stack (Loki, Grafana, Tempo, Mimir) on Kubernetes
Instrument applications in any language using OpenTelemetry SDKs and configure the OpenTelemetry Collector as an enterprise telemetry pipeline
Build production-grade Prometheus monitoring with PromQL expertise, recording rules, AlertManager routing, and Thanos for long-term scalability
Master Grafana deeply — advanced dashboard engineering, templating, alerting, Grafana OnCall, and dashboards-as-code with Grafonnet and Jsonnet
Implement and operate Datadog as a complete enterprise observability platform — APM, logs, metrics, RUM, Synthetics, and monitors-as-code with Terraform
Configure and operate Dynatrace for AI-powered full-stack observability and anomaly detection
Build and manage the ELK Stack (Elasticsearch, Logstash, Kibana) for enterprise log management at scale
Instrument distributed systems with distributed tracing — understanding trace context propagation, sampling strategies, and tail-based sampling
Build comprehensive AWS CloudWatch observability — custom metrics, Container Insights, X-Ray, Synthetics, and cross-account monitoring
Implement AIOps using ML-based anomaly detection, intelligent alerting, and AI-assisted root cause analysis
Design observability cost governance strategies — cardinality reduction, log filtering, sampling, and telemetry pipeline optimisation
Build everything as code — dashboards, alerts, SLOs, and monitoring configurations defined in Git and deployed via CI/CD
Course Curriculum
9 Sections · 0 LessonsOpenTelemetry — The Universal Instrumentation Standard
Prometheus — Metrics at Scale
Grafana — Dashboard Engineering & Unified Observability
Log Management — Loki, ELK Stack & Enterprise Log Analytics
Distributed Tracing — End-to-End Request Visibility
Commercial Observability Platforms — Datadog, Dynatrace & New Relic
AWS Native Observability — CloudWatch, X-Ray & Container Insights
AIOps, Continuous Profiling & Emerging Observability
Observability at Scale — Cost, Governance & Platform Engineering
Ratings & Reviews
0 reviews