⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 182 of 186 in AWS & Cloud Architecture
Senior DevOps / SRE Multi-Cloud Multi-Cloud Observability & SRE Observability

Q: A customer transaction starts at an Azure Front Door gateway, calls an Azure Function, triggers a microservice on Google Kubernetes Engine (GKE), and updates a DynamoDB database in AWS. When a request takes 4 seconds instead of 200ms, how do you trace the request end-to-end across all three clouds using OpenTelemetry?

Engineering a vendor-neutral distributed tracing architecture across AWS Lambda, GKE microservices, and Azure Functions using the OpenTelemetry Collector and Grafana Tempo.

#Multi-Cloud #OpenTelemetry #Distributed Tracing #Grafana Tempo #AWS #GCP #Azure
🎙️ Candidate Opening & Architectural Context
"Cloud-specific tracing tools (AWS X-Ray, Google Cloud Trace, Azure Application Insights) cannot correlate spans across cloud boundaries, leaving black holes in transaction lifecycles. We built a unified multi-cloud tracing plane using OpenTelemetry standards."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Enforce W3C Trace Context Standard Across Application runtimes

Standardize trace propagation headers across distributed microservices:

  • W3C Header Standard: Configured OpenTelemetry SDKs in Node.js, Go, and Python to propagate traceparent and tracestate HTTP headers.
  • Context Propagation: When Azure Function invokes the GKE microservice, it passes traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01.
Pro Tip: W3C Trace Context is an open industry standard that ensures trace identity is preserved across heterogeneous clouds, runtimes, and message brokers.
2️⃣

Deploy OpenTelemetry Collectors in Each Cloud Environment

Collect, buffer, and batch telemetry locally before transmitting across networks:

  • Deployment Topologies: Deployed OTel Collector as a DaemonSet in GKE, an AWS Lambda Layer in AWS, and an Azure Container App in Azure.
  • Batch & Compression: Configured batch processor with send_batch_max_size: 1000 and timeout: 5s, compressing spans with gzip over gRPC (OTLP/gRPC).
Pro Tip: Running local OpenTelemetry Collectors in each cloud decouples application code from tracing storage backends and buffers telemetry during network interruptions.
3️⃣

Consolidate Traces in Centralized Grafana Tempo Storage

Store petabytes of trace data cost-effectively in cloud object storage:

  • Tempo Cluster: Deployed distributed Grafana Tempo backed by cost-efficient S3/GCS object storage.
  • Trace Search: Enabled TraceQL queries allowing engineers to filter spans by cloud provider, HTTP status code, and latency duration: { .cloud.provider = 'gcp' && duration > 2s }.
Pro Tip: Grafana Tempo stores traces directly in object storage without indexing every span attribute, reducing tracing storage costs by 80% compared to Elasticsearch.
4️⃣

Pinpoint Cross-Cloud Latency Bottlenecks in Grafana Dashboards

Visualize end-to-end waterfall flame graphs spanning multi-cloud requests:

  • Waterfall Trace: Visualized the exact 4-second request in Grafana: 45ms in Azure Function, 3,820ms waiting for a slow GCP Cloud SQL query, and 60ms in AWS DynamoDB.
  • Root Cause Resolution: Optimized the missing index in GCP Cloud SQL, immediately reducing total transaction latency back to 140ms.
Pro Tip: End-to-end tracing replaces inter-team finger-pointing with definitive millisecond-level breakdown evidence.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"OpenTelemetry with W3C Trace Context standardizes distributed tracing across AWS, GCP, and Azure, enabling end-to-end request visibility in Grafana Tempo without cloud vendor lock-in."
⚡ 60-Second Elevator Pitch Talking Points
  • Enforce W3C Trace Context HTTP header propagation across all services in all clouds.
  • Deploy OpenTelemetry Collectors locally in AWS, GCP, and Azure to buffer and batch telemetry.
  • Stream OTLP/gRPC spans into a centralized, object-storage-backed Grafana Tempo cluster.
  • Use TraceQL and waterfall flame graphs to pinpoint exact cross-cloud latency bottlenecks.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →