Q: A customer transaction starts at an Azure Front Door gateway, calls an Azure Function, triggers a microservice on Google Kubernetes Engine (GKE), and updates a DynamoDB database in AWS. When a request takes 4 seconds instead of 200ms, how do you trace the request end-to-end across all three clouds using OpenTelemetry?
Engineering a vendor-neutral distributed tracing architecture across AWS Lambda, GKE microservices, and Azure Functions using the OpenTelemetry Collector and Grafana Tempo.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Enforce W3C Trace Context Standard Across Application runtimes
Standardize trace propagation headers across distributed microservices:
- W3C Header Standard: Configured OpenTelemetry SDKs in Node.js, Go, and Python to propagate
traceparentandtracestateHTTP headers. - Context Propagation: When Azure Function invokes the GKE microservice, it passes
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01.
Deploy OpenTelemetry Collectors in Each Cloud Environment
Collect, buffer, and batch telemetry locally before transmitting across networks:
- Deployment Topologies: Deployed OTel Collector as a DaemonSet in GKE, an AWS Lambda Layer in AWS, and an Azure Container App in Azure.
- Batch & Compression: Configured
batchprocessor withsend_batch_max_size: 1000andtimeout: 5s, compressing spans with gzip over gRPC (OTLP/gRPC).
Consolidate Traces in Centralized Grafana Tempo Storage
Store petabytes of trace data cost-effectively in cloud object storage:
- Tempo Cluster: Deployed distributed Grafana Tempo backed by cost-efficient S3/GCS object storage.
- Trace Search: Enabled TraceQL queries allowing engineers to filter spans by cloud provider, HTTP status code, and latency duration:
{ .cloud.provider = 'gcp' && duration > 2s }.
Pinpoint Cross-Cloud Latency Bottlenecks in Grafana Dashboards
Visualize end-to-end waterfall flame graphs spanning multi-cloud requests:
- Waterfall Trace: Visualized the exact 4-second request in Grafana: 45ms in Azure Function, 3,820ms waiting for a slow GCP Cloud SQL query, and 60ms in AWS DynamoDB.
- Root Cause Resolution: Optimized the missing index in GCP Cloud SQL, immediately reducing total transaction latency back to 140ms.
- Enforce W3C Trace Context HTTP header propagation across all services in all clouds.
- Deploy OpenTelemetry Collectors locally in AWS, GCP, and Azure to buffer and batch telemetry.
- Stream OTLP/gRPC spans into a centralized, object-storage-backed Grafana Tempo cluster.
- Use TraceQL and waterfall flame graphs to pinpoint exact cross-cloud latency bottlenecks.