Q: What is the difference between metrics, logs, and traces, and in what order do you triage them during an incident?
Systematic incident triaging progression across the three pillars of observability: starting with Metrics to detect and scope the blast radius, pivoting to Distributed Tracing to pinpoint the failing service hop, and drilling into Logs for root-cause stack traces.
#Observability #Metrics #Logs #Traces #OpenTelemetry #Prometheus #Jaeger #SRE
🎙️ Candidate Opening & Architectural Context
"Metrics are aggregatable numeric time-series data ideal for real-time alerting and scoping trends. Logs are discrete, structured event records providing deep textual context for specific operations. Distributed Traces track the end-to-end journey of a request across distributed microservice hops, isolating network latency and dependency failures. During an incident, I triage them in a disciplined order: Metrics first to detect and scope the blast radius, Traces second to isolate the bottleneck hop, and Logs third to inspect the exact exception and root cause."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Comparative Roles of the Three Telemetry Pillars
Understanding the unique strength and limitation of each signal:
# Step 1: Metric check with Prometheus (detect spike in 5xx errors)
rate(http_requests_total{job="checkout-service", status=~"5.."}[5m])
# Calculate p99 latency by endpoint
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{job="checkout-service"}[5m])) by (le, path))
- Metrics (Detection & Scoping): Answer 'Is there an issue? How big is it? When did it start? Which service/region/tenant is affected?' Metrics are cheap to store, have low latency, and drive our alert rules.
- Distributed Traces (Pinpointing & Latency): Answer 'Where in the call graph is the request spending time or failing?' Tracing breaks down p99 latency across 15 microservices down to the millisecond.
- Logs (Root Cause & Context): Answer 'Why did this specific operation fail?' Logs contain the stack trace, error messages, user payload IDs, and exact line numbers.
2️⃣
The Step-by-Step Incident Triage Progression
How SREs navigate from high-level alerts to the exact bug in under 5 minutes:
# Correlate trace_id into Kubernetes logs directly
kubectl logs -l app=payment-service -n prod --since=10m | grep 'trace_id=4bf92f3577b34da6a3ce929d0e0e4736'
# Look for specific connection timeouts
kubectl logs deploy/payment-service -n prod --since=15m | grep -iE 'timeout|connection refused'
- 1. Metrics First: An alert fires. Check Grafana dashboards to identify the affected service, traffic volume, and blast radius (e.g., only in us-west-2 or affecting all tenants).
- 2. Traces Second: Inspect trace spans in Jaeger/Tempo. Identify that checkout-service latency is normal, but it is waiting 10 seconds on the payment-gateway HTTP hop.
- 3. Logs Third: Grab the unique
trace_idfrom the slow span, filter Loki/Datadog logs for that exacttrace_id, and read the exception:SocketTimeoutException: Connection reset by peer.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never dive into logs first during an outage — millions of log lines create noise and tunnel vision. Follow the proven SRE progression: Metrics to scope -> Traces to isolate the hop -> Logs with correlation IDs for root cause."
⚡ 60-Second Elevator Pitch Talking Points
- Triage systematically: Metrics to detect scope, Traces to pinpoint the failing microservice hop, and Logs for the exact stack trace.
- Avoid log diving early in an incident to prevent getting overwhelmed by misleading noise.
- Enforce OpenTelemetry correlation IDs so engineers can pivot seamlessly from a high-latency trace span directly to the matching log line.
Advertisement