AZ-DEV-190: Observability with Application Insights and OpenTelemetry
About This Course
Observability separates the applications that get diagnosed in minutes from the ones that get diagnosed in days. Learn the modern OpenTelemetry-first approach for .NET 10 on Azure: instrument traces, metrics, and logs; write KQL to answer real production questions; configure SLO-driven alerts; author workbooks; use Live Metrics and Profiler for real-time diagnosis; correlate distributed traces across queues, functions, and external services; and drive telemetry cost down without losing signal. By the end you will be able to design, ship, and operate a full-stack observability program for a multi-service .NET application on Azure.
Course Curriculum
20 Lessons
AZ-DEV-190 M1L1 - Observability fundamentals - three pillars and the OpenTelemetry model
The three pillars of observability (metrics, logs, traces), the OpenTelemetry model (signals, resources, attributes, propagation), Application Insights architecture (workspace-based, KQL surface), and comparison to Prometheus + Grafana + Loki.
AZ-DEV-190 M1L2 - Instrument a .NET 10 minimal API with OpenTelemetry to Azure Monitor - Lab Exercises
Note: This lab pre-provisions an empty resource group at start — allow up to 3 minutes for it to become ready before beginning the exercises.
Create Anchorline's observability foundation yourself with the Azure CLI — a Log Analytics workspace, workspace-based Application Insights, a Linux App Service Plan, and a Web App — then wire the site to App Insights via APPLICATIONINSIGHTS_CONNECTION_STRING. Deploy the .NET 10 minimal API with Azure.Monitor.OpenTelemetry.AspNetCore (the Distro), push traffic, and verify traces, metrics, and logs land in App Insights.
AZ-DEV-190 M2L3 - Instrumenting ASP.NET Core with automatic and manual OpenTelemetry
Automatic instrumentation (AspNetCore, HttpClient, SqlClient), manual instrumentation with ActivitySource and Meter, baggage propagation, W3C Trace Context.
AZ-DEV-190 M2L4 - Trace a 3-service .NET solution end-to-end with baggage propagation - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Instrument a 3-service .NET 10 solution (Web → API → Worker). Verify a single trace spans all three services and shows the SQL calls at the leaves. Add baggage carrying TenantId through the call chain and see it appear on every span.
AZ-DEV-190 M3L5 - Custom telemetry - activities, counters, histograms, and structured logs
ActivitySource for custom spans, Meter + Counter<T> / Histogram<T> / ObservableGauge<T> for custom metrics, ILogger with OTel export, and the mental model for SetAttribute vs AddEvent vs SetStatus.
AZ-DEV-190 M3L6 - Emit business-domain telemetry in the Orders service - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Add domain telemetry: orders.received.total counter, orders.processing.duration.ms histogram, Orders.Validate custom activity with orderId, customerId, totalAmount attributes. Verify each in App Insights.
AZ-DEV-190 M4L7 - KQL for developers - troubleshooting Azure Monitor telemetry
KQL fundamentals — where, project, extend, summarize, join, let; time-window queries with ago and bin; common troubleshooting patterns; performance idioms.
AZ-DEV-190 M4L8 - Answer 10 troubleshooting questions with KQL - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Answer 10 troubleshooting questions with KQL against live App Insights telemetry: slowest endpoints, error rates, correlated errors between services, dependency-failure breakdowns, and p50/p95/p99 latency per operation.
AZ-DEV-190 M5L9 - Alerts and action groups - SLOs, burn rates, and notification routing
Metric alerts vs log alerts, dynamic vs static thresholds, action groups (email, webhook, Function, Logic App, ITSM), SLO-based alerts using burn rates, alert suppression, alert-fatigue design.
AZ-DEV-190 M5L10 - Configure fast-burn and slow-burn SLO alerts wired to Teams and Logic Apps - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Define an SLO (99.5% of Orders API requests <500ms), configure fast-burn (2% budget in 1h) and slow-burn (5% in 6h) alerts, route to an action group with a Teams webhook and Logic App runbook. Trigger by induced latency and verify.
AZ-DEV-190 M6L11 - Workbooks and dashboards for developer self-service ops
Workbook architecture — parameters, queries, visualizations, drill-down. Shared vs private workbooks, workbook templates, and how they compare to Grafana dashboards.
AZ-DEV-190 M6L12 - Ship a service-health workbook with drill-down to end-to-end trace - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Build a workbook for the Orders API showing request rate, error rate, p95 latency broken down by revision with a time-range parameter, plus drill-down from any single request into its full distributed trace. Publish for the team.
AZ-DEV-190 M7L13 - Live Metrics, Profiler, and Snapshot Debugger for real-time diagnosis
Live Metrics Stream (real-time, no query lag), Application Insights Profiler (production sampling profiler), Snapshot Debugger for exception snapshots, and when each is worth reaching for.
AZ-DEV-190 M7L14 - Diagnose a hot path in production with Profiler and fix it - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Enable Profiler on the Orders API, trigger a memory-heavy operation (large JSON serialization), analyze the Profiler snapshot to identify the hot path, and fix the root cause (JsonSerializer reuse). Confirm the improvement in Live Metrics.
AZ-DEV-190 M8L15 - Distributed tracing at scale - propagation across HTTP, queues, and functions
W3C Trace Context propagation across HTTP, Service Bus, Event Hubs, Storage Queues. Correlating a trace across HTTP → SB queue → Function → Cosmos write → downstream HTTP. Trace context in async workflows. Diagnostic.Activity.Current.
AZ-DEV-190 M8L16 - Trace a message end-to-end across HTTP, queue, function, Cosmos - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 12 minutes for the environment to become ready before beginning the exercises.
Trace a message that flows HTTP intake → Service Bus queue → Azure Function → Cosmos write → downstream HTTP call. Verify a single trace ID surfaces every hop in the App Insights end-to-end trace view.
AZ-DEV-190 M9L17 - Telemetry cost optimization - sampling, retention, and workspace design
Ingestion pricing model, fixed vs adaptive sampling, Log Analytics workspace design (basic vs analytics logs, retention tiers), commitment tiers, ingestion caps, and query-only tables.
AZ-DEV-190 M9L18 - Cut telemetry ingestion cost 60% with adaptive sampling and tier routing - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 10 minutes for the environment to become ready before beginning the exercises.
Take a service ingesting simulated 100 GB/day. Apply adaptive sampling (target 5%) via OTel TraceIdRatioBasedSampler. Move debug-level logs to a basic-logs table. Configure retention tiers. Measure the resulting cost reduction — target 60%+ cut.
AZ-DEV-190 M10L19 - Capstone - Observability program design with SLOs, error budgets, and runbooks
Observability program design — SLOs, SLIs, error budgets, runbook automation, postmortem culture, cost vs coverage tradeoffs, on-call rotation as a design input.
AZ-DEV-190 M10L20 - Capstone - Ship full-stack observability for a 4-service Anchorline system - Lab Exercises
Note: This lab pre-provisions Azure resources at start — allow up to 15 minutes for the environment to become ready before beginning the exercises.
Instrument a 4-service .NET 10 Anchorline system (Web + Orders API + Fulfillment Worker + Notification Function) end-to-end with OpenTelemetry + App Insights. Define 3 SLOs. Build a service-catalog workbook. Configure fast + slow burn alerts to Teams + a Logic App runbook. Inject latency into Cosmos calls and verify the on-call runbook completes root-cause analysis in under 5 minutes.