Q: How would you monitor end-to-end SLA for services involved in a payments pipeline?
Design an end-to-end SLA and error budget monitoring architecture for a mission-critical multi-hop payments processing pipeline.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Define Composite Multi-Hop Service Level Indicators (SLIs)
Establish the core SLIs for the payments pipeline: - **Availability SLI**: `Successful Payments (HTTP 200 / Approved) / Total Initiated Payments` >= 99.99%. - **Latency SLI**: `p99 End-to-End Payment Duration (from client request to bank confirmation)` < 1,500ms. - **Freshness SLI**: Asynchronous settlement event queue lag < 5 seconds.
Propagate Distributed Trace Context via OpenTelemetry (W3C TraceContext)
Enforce W3C `traceparent` context propagation across all microservices and asynchronous message queues (Kafka / SQS). Every payment carries a unique `payment_id` and `trace_id` span attribute, enabling instant lookup of any degraded transaction.
# OpenTelemetry baggage propagation in HTTP headers
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
tracestate: payment_id=pay_998124,tier=enterprise
Multi-Window Multi-Burn-Rate Alerting on Error Budgets
Implement Google SRE multi-window burn-rate alerts in Prometheus / Alertmanager. Trigger page alerts only when the rate of error budget consumption threatens the monthly SLA (e.g. burning 2% of the budget in 1 hour), eliminating false alarm pager fatigue.
# PromQL Multi-Burn-Rate Alert Rule
expr: |
(sum(rate(payment_failures_total[1h])) / sum(rate(payment_attempts_total[1h])) > (14.4 * 0.0001))
and
(sum(rate(payment_failures_total[5m])) / sum(rate(payment_attempts_total[5m])) > (14.4 * 0.0001))
for: 2m
labels:
severity: critical
team: payments-sre
Deploy Synthetic Payment Probes in Production
Execute continuous synthetic end-to-end transactions every 60 seconds using test credit card numbers against internal banking mock targets to detect pipeline degradation before real customers are impacted.
- Define composite availability and latency SLIs spanning the entire multi-hop payment lifecycle.
- Enforce OpenTelemetry W3C trace context propagation across HTTP services and Kafka queues.
- Implement multi-window burn-rate alerting to track monthly error budgets and prevent alert fatigue.
- Run continuous synthetic transactions to detect pipeline regressions before customers do.