Design an Observability Strategy for Distributed Production Systems
Full-stack observability architecture: Google's Four Golden Signals, OpenTelemetry standard, log-metric-trace correlation with exemplars, multi-window erro...
Master 85+ battle-tested scenario-based Observability & Monitoring interview questions for Senior DevOps, Cloud, and SRE engineers. Includes incident runbooks, STAR talking points, and CLI commands.
Full-stack observability architecture: Google's Four Golden Signals, OpenTelemetry standard, log-metric-trace correlation with exemplars, multi-window erro...
If the application instance's local resources are fine, the latency is almost certainly caused by an external downstream dependency....
Waking up for non-actionable, self-resolving alerts creates alert fatigue and burns out engineers. The alert is poorly designed for this ......
Logs do not magically appear in aggregators; they traverse a pipeline. I would check:...
Prometheus OOMs almost exclusively due to High Cardinality in the metrics it scrapes or the queries being run against it....
1. Metrics: Time-series aggregated numbers (e.g., requests_per_second, cpu_usage). They are cheap to store and allow for fast alerting an......
Availability shouldn't be measured purely by ping or CPU (host uptime), because the host could be up but the app returning 500 errors. ...
This is solved using Trace IDs or Correlation IDs....
Querying raw data over long periods (e.g., aggregating 1 year of CPU data on the fly) involves analyzing billions of data points, choking......
A static threshold often fails because it ignores the rate of change. ...
Elasticsearch cluster states are:...
SRE uses Error Budgets to make data-driven decisions between feature velocity and reliability, removing the emotion from the conversation....
1. Push: The application (or an agent on the host) writes metrics actively and sends them over the network to a centralized aggregator en......
If the IAM permissions are correct (i.e., logs:CreateLogStream, logs:PutLogEvents), the issue is often configuration:...
Distributed tracing is computationally expensive and memory-intensive because it tracks every span of a request. You should never trace 1......
Google's SRE book defines four "Golden Signals" as the baseline for user-facing systems:...
Inside less I would:...
Alerting on the age of the oldest item is significantly better....
SRE teams must implement a multi-tier storage architecture, often called Hot/Warm/Cold tiers....
The missing protection mechanism is a Circuit Breaker combined with Timeouts....
1. Counter: A cumulative metric that can only go up (or reset to zero on restart). Examples include http_requests_total or bytes_sent. Be......
- Whitebox Monitoring: Depends on the internal state and telemetry exposed by the system itself (e.g., APM, custom app metrics, logs). It......
- SLI (Service Level Indicator): A quantitative, mathematical measure of some aspect of the level of service provided. Example: "The perc......
An Average hides extreme outliers. If 99 users experience lightning-fast 10ms latencies, but 1 user hits a database timeout and waits 5,0......
Standard HTTP tracing relies on passing headers (like traceparent). A message queue breaks the HTTP chain....
Standard infrastructural monitoring only cares about HTTP status codes. To catch semantic/logic errors from third parties, you must imple......
Synthetic Monitoring is simulating user traffic to proactively test your systems from the outside....
This is called a Flapping Alert. To fix it, you introduce Hysteresis or a Pending Duration....
You use the Prometheus Pushgateway....
Modern cloud-native applications must support Dynamic Log Level adjustments without process restarts....
When you cannot modify the application code (zero-instrumentation), you use eBPF-based observability....
Alerting on a static threshold (e.g., "Alert if 100 errors happen") is flawed because it ignores traffic volume: 100 errors out of 100 re......
The Apdex score is an open standard to translate raw latency numbers into a single metric representing global user satisfaction, ranging ......
Metrics are highly aggregated (e.g., "You had 50 requests take longer than 2 seconds"). Traces are highly specific. The painful gap histo......
If Team A tags their database queries as {"db.table_name": "users"}, Team B uses {"db_table": "users"}, and Team C uses ...
A Dead Letter Queue (DLQ) is a secondary queue where an asynchronous system routes messages that completely fail to be processed after mu......
High TSDB compaction (and resulting IOPS thrashing) is heavily correlated with Metric Churn....
Backend APM measures performance from the moment the request hits your data center's load balancer until the server finishes processing it....
I would implement Continuous Profiling....
- Liveness Probe: Checks if the application container is fundamentally healthy and running. If the liveness probe fails (e.g., the app is......
Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion....
Average latency is deceptive because it hides the temporal distribution of requests. A system can have a "good" average while users exper......
This is alert flapping—when a metric oscillates around the threshold, causing rapid alert cycles. Engineers ignore the notifications (ale......
Each unique combination of label values creates a separate time-series. Prometheus stores each series' metadata, recent data points, and ......
Symptoms are resource metrics (CPU, memory, disk). Root causes are user-facing impacts (errors, latency, requests failing)....
Log sampling reduces volume while preserving critical signals. There are two approaches:...
Cost reduction requires architectural restructuring—not just turning off features. The strategy is Hot/Warm/Cold tiers with intelligent r......
A good runbook is not a novel—it's a decision tree. It guides humans through uncertainty without requiring them to read 50 pages at 3 AM....
If spans are created inside the application but do not reach the backend, the problem is usually between the SDK and the telemetry backend....
I would treat the collector as production infrastructure, not a side experiment....
A Histogram stores observations in configurable buckets, such as requests under 100ms, 300ms, 1s, and 5s. Prometheus can aggregate histog......
For a classic Prometheus histogram, I would calculate P95 from the bucket rate:...
Classic histograms multiply time series by every bucket and every label combination. If a metric has 20 buckets and many labels, cost and......
This is a dependency fan-out problem. The database is the likely root cause, while the application alerts are symptoms....
A health check should be cheap, fast, and designed for the action that will be taken when it fails....
Dashboards should show deploy context and user impact, not just raw spikes....
The services are using different propagation formats. One service injects trace context using headers such as traceparent and tracestate,......
I would fix this at multiple layers because relying on one filter is risky....
Structured logging writes logs as key-value data, usually JSON:...
Missing data does not mean zero. It usually means Prometheus did not scrape the target, the metric disappeared, or the pod restarted and ......
Multi-window burn-rate alerts combine a short window and a long window....
Aggregate dashboards hide small-scope failures. If the canary has a 20% error rate but receives only 5% of traffic, the global error rate......
Monitoring tells you whether known failure modes are happening. Example: CPU is high, disk is full, or the service is returning 500 errors....
The synthetic checks do not match the user population or the failing path. A single US probe cannot prove global availability....
Service mesh telemetry is powerful but can create a series for every source-destination-path combination....
I would design logging so logs leave the process quickly and survive container restarts....
Regular gaps point to scrape or ingestion problems, not random application behavior....
I would combine tail-based sampling with strict attribute controls....
MTTD means Mean Time To Detect: how long it takes to notice a problem after it starts....
Severity should map to required human action....
Serverless observability must use the platform telemetry path because functions may start and disappear before a pull-based scraper can r......
Host CPU and memory are not the only bottlenecks. I would check saturation inside the application and dependencies:...
An absence alert fires when expected telemetry stops arriving....
Backend APM cannot see browser rendering failures....
Observability systems need their own monitoring....
The most common cause is clock skew between hosts, containers, or regions. Distributed systems rely on timestamps for log ordering, trace......
Without consistent metadata, telemetry is hard to search and easy to misread....
I would troubleshoot the scrape path from Prometheus to the pods....
Remote storage must be treated as a dependency that can fail....
The span name contains unbounded identifiers. Every user ID and order ID creates a different operation name, making search, aggregation, ......
I would build the dashboard around the customer journey, not servers....
End-to-end diagnostic RCA runbook for troubleshooting a silent mTLS handshake failure between an edge Ingress Gateway and internal mesh sidecars following a con...
Master production incident triage for a distributed streaming failure where Cloud Load Balancers report healthy target groups, mesh sidecars show green, yet end...
How to design pre-launch chaos experiments, synthetic stress tests, and automated tiered graceful degradation for a global high-concurrency streaming premiere....
Systematic incident triaging progression across the three pillars of observability: starting with Metrics to detect and scope the blast radius, pivoting to Dist...