Free through lesson 3

Monitoring & Observability for Beginners

From telling symptoms apart from causes to structured logs, Prometheus metrics, Grafana dashboards, Alertmanager notifications, OpenTelemetry distributed tracing, and SLOs with error budgets. Over 30 lessons, you'll get to the point where you notice problems in your own service before your users do. You'll actually run Prometheus, Grafana, Alertmanager, and OpenTelemetry on your own machine, and every output shown here was really measured.

Curriculum

The 30 lessons are split into 6 chapters. We recommend going in order from chapter 1, but feel free to jump to whatever interests you. Note: commands can't run in the browser, so try them on your own machine with Prometheus, Grafana, and OpenTelemetry installed. The sample app runs on the Python standard library alone. Domains and host names have been replaced with example values (such as shop.example.com).

Chapter 1 — What Do You Want to Know About? (lessons 1–4)

Before monitoring is about tools, it's about what you want to know. You'll separate symptoms from causes, pick the four signals worth measuring, and decide up front when it's worth waking someone up.

Chapter 2 — Logs (lessons 5–11)

First, you'll make your logs searchable after the fact. You'll structure them, settle on levels, strip out personal data, and get to the point where you can follow a single request with a trace_id.

Chapter 3 — Metrics (lessons 12–18)

Keeping numbers as time series reveals trends and anomalies. You'll write your own /metrics, have Prometheus come and collect it, and use PromQL to get latency and error rates.

Chapter 4 — Visualization and Alerts (lessons 19–23)

You'll turn numbers into something people can look at, and when nobody is looking, call them with a notification. You'll build Grafana dashboards and set up Alertmanager receivers, grouping, and silences.

Chapter 5 — Distributed Tracing (lessons 24–27)

You'll see where a single request spent how many milliseconds. You'll emit spans with OpenTelemetry, connect them across services, and learn to name the slow path.

Chapter 6 — Keeping It Going (lessons 28–30)

Monitoring isn't done once it's built. You'll use SLOs and error budgets to decide between "fix things or ship features," and keep it all running with on-call and postmortems.

Once you've finished all 30 lessons, move on to GitHub Actions Fundamentals Course, which automates everything from push to deploy, or Terraform Fundamentals Course, which sets up the servers themselves as code. If you're still unsure about the basics of deploying, go back to Intro to Shipping and Running a Service. A membership unlocks every course.