Q: How do you design a centralized multi-tenant log aggregation architecture using OpenTelemetry Collector and Grafana Loki that guarantees namespace isolation, prevents log ingestion runaway costs, and scales across multiple Kubernetes clusters?
Designing and operating an enterprise multi-tenant logging pipeline utilizing OpenTelemetry Collector DaemonSets and Grafana Loki with strict tenant isolation, rate limiting, and cost controls.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Edge Collection and Metadata Enrichment with OpenTelemetry Collector
Deploy OpenTelemetry Collector as a DaemonSet with filelog receiver. Extract Kubernetes metadata (namespace, pod, container, git_commit) using the k8sattributes processor. Sanitize logs to mask PII (credit cards, JWTs, API tokens) using transformprocessor before logs ever leave the node.
processors:
k8sattributes:
auth_type: 'serviceAccount'
extract:
metadata: [k8s.namespace.name, k8s.pod.name, k8s.container.name]
transform:
log_statements:
- context: log
statements:
- replace_pattern(body, '(?i)(bearer\s+)[a-zA-Z0-9_\-\.]+', '$$1[REDACTED]')
Multi-Tenant Routing and Header Injection
Dynamically route logs into Grafana Loki using the Loki exporter, injecting X-Scope-OrgID headers mapped to Kubernetes namespaces or business units. This enables Loki's native microservices-mode multi-tenancy, isolating log search queries and retention policies between teams.
exporters:
loki:
endpoint: 'http://loki-gateway.logging:80/loki/api/v1/push'
headers:
'X-Scope-OrgID': 'tenant-team-a'
FinOps Guardrails and Dynamic Ingestion Rate Limiting
Configure Loki overrides and tenant limits (ingestion_rate_mb, max_streams_per_user, burst_size_mb). Implement loki-canary and Prometheus alerting on tenant throughput to automatically drop debug-level chatter or throttle noisy microservices before network and storage degradation occurs.
limits_config:
per_tenant_override_config: /etc/loki/overrides.yaml
ingestion_rate_mb: 10
ingestion_burst_size_mb: 20
max_streams_per_user: 10000
- We run OpenTelemetry Collector as a node DaemonSet to scrub PII and inject Kubernetes metadata at the edge.
- Logs are routed to Grafana Loki in microservices mode with dynamic `X-Scope-OrgID` tenant headers mapping to namespaces.
- Per-tenant ingestion limits and S3 object storage slashed logging costs by 70% compared to legacy Elasticsearch clusters.