⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Platform Engineering & IDP Interview Questions Scenario 41 of 50 in Platform Engineering & IDP
Staff Platform Engineer Platform Engineering Observability & FinOps Observability
🎯 Target Role / Context: Staff Platform Engineer building telemetry infrastructure and observability platforms.

Q: How do you design a centralized multi-tenant log aggregation architecture using OpenTelemetry Collector and Grafana Loki that guarantees namespace isolation, prevents log ingestion runaway costs, and scales across multiple Kubernetes clusters?

Designing and operating an enterprise multi-tenant logging pipeline utilizing OpenTelemetry Collector DaemonSets and Grafana Loki with strict tenant isolation, rate limiting, and cost controls.

#Loki #OpenTelemetry #Logging #Multi-Tenancy #FinOps #Kubernetes
🎙️ Candidate Opening & Architectural Context
"Uncontrolled log streaming quickly saturates storage systems and creates massive cloud billing surprises. A robust platform observability tier must ingest container stdout logs, parse metadata at the edge via OpenTelemetry Collector DaemonSets, enforce tenant quotas, and store chunked indexes economically in object storage using Grafana Loki."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Edge Collection and Metadata Enrichment with OpenTelemetry Collector

Deploy OpenTelemetry Collector as a DaemonSet with filelog receiver. Extract Kubernetes metadata (namespace, pod, container, git_commit) using the k8sattributes processor. Sanitize logs to mask PII (credit cards, JWTs, API tokens) using transformprocessor before logs ever leave the node.

processors:
  k8sattributes:
    auth_type: 'serviceAccount'
    extract:
      metadata: [k8s.namespace.name, k8s.pod.name, k8s.container.name]
  transform:
    log_statements:
      - context: log
        statements:
          - replace_pattern(body, '(?i)(bearer\s+)[a-zA-Z0-9_\-\.]+', '$$1[REDACTED]')
2

Multi-Tenant Routing and Header Injection

Dynamically route logs into Grafana Loki using the Loki exporter, injecting X-Scope-OrgID headers mapped to Kubernetes namespaces or business units. This enables Loki's native microservices-mode multi-tenancy, isolating log search queries and retention policies between teams.

exporters:
  loki:
    endpoint: 'http://loki-gateway.logging:80/loki/api/v1/push'
    headers:
      'X-Scope-OrgID': 'tenant-team-a'
Advertisement
3

FinOps Guardrails and Dynamic Ingestion Rate Limiting

Configure Loki overrides and tenant limits (ingestion_rate_mb, max_streams_per_user, burst_size_mb). Implement loki-canary and Prometheus alerting on tenant throughput to automatically drop debug-level chatter or throttle noisy microservices before network and storage degradation occurs.

limits_config:
  per_tenant_override_config: /etc/loki/overrides.yaml
  ingestion_rate_mb: 10
  ingestion_burst_size_mb: 20
  max_streams_per_user: 10000
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Multi-tenant logging requires edge filtering with OpenTelemetry Collector to scrub PII and append Kubernetes metadata, paired with Loki's multi-tenancy headers (X-Scope-OrgID) and tenant rate limits to prevent cost blowouts and maintain query isolation."
⚡ 60-Second Elevator Pitch Talking Points
  • We run OpenTelemetry Collector as a node DaemonSet to scrub PII and inject Kubernetes metadata at the edge.
  • Logs are routed to Grafana Loki in microservices mode with dynamic `X-Scope-OrgID` tenant headers mapping to namespaces.
  • Per-tenant ingestion limits and S3 object storage slashed logging costs by 70% compared to legacy Elasticsearch clusters.
Advertisement
Want more Platform Engineering & IDP scenarios?
Explore our complete collection of scenario-based Platform Engineering & IDP interview runbooks.
Browse All Platform Engineering & IDP Questions →