⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Terraform & IaC Interview Questions Scenario 111 of 117 in Terraform & IaC
Senior DevOps / Platform Engineer Terraform Drift Detection & State Reconciliation Critical IaC Triage

Q: A terraform apply is failing due to configuration drift, but the underlying infrastructure is live and mission-critical. How do you reconcile state and fix the drift without causing downtime?

Production runbook for safely reconciling configuration drift during a failing terraform apply on mission-critical live infrastructure without risking service interruption.

#Terraform #IaC #Configuration Drift #State Management #terraform refresh #Zero Downtime
🎙️ Candidate Opening & Architectural Context
"Configuration drift occurs when live cloud resources are modified outside of Terraform (via cloud console, automated autoscalers, or emergency hotfixes), causing Terraform's expected state to diverge from reality. When `terraform apply` fails or threatens destructive replacement on critical systems, you must never force apply. You must isolate the drift, refresh the state safely, backport modifications, and protect against unintended destruction."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's HashiCorp Certified Terraform Associate (003) Interactive Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Run Non-Destructive Speculative Plan & Backup Remote State

Immediately create a backup snapshot of the current remote state file before running any write operations. Run `terraform plan` and analyze every resource flagged with `~` (update in-place), `+` (create), and especially `-` or `- / +` (destroy and recreate).

# Backup remote state
terraform state pull > terraform-state-backup-$(date +%s).json
# Run targeted speculative plan without touching live infrastructure
terraform plan -detailed-exitcode -out=tfplan.binary
terraform show -json tfplan.binary > plan_analysis.json
2

Isolate Why the Resource Threatens Destruction

If Terraform plans to destroy and recreate a live resource, identify the argument forcing recreation (e.g. changing an EC2 AMI, RDS storage type, or VPC subnet CIDR). Add `lifecycle { create_before_destroy = true }` or `lifecycle { ignore_changes = [ ... ] }` in your HCL code to block destructive replacements.

# Prevent accidental destruction of live databases/clusters
lifecycle {
  prevent_destroy = true
  ignore_changes = [
    tags["LastModified"],
    allocated_storage,
    security_groups
  ]
}
Advertisement
3

Refresh State from Reality and Backport Manual Changes into HCL

Run `terraform apply -refresh-only` (in Terraform 1.1+) to update Terraform's state file with real-world infrastructure attributes without applying configuration changes. Once state matches cloud reality, inspect the diff and update your Terraform HCL code to mirror the live configuration.

# Inspect and approve state refresh safely
terraform apply -refresh-only
# Update HCL code to match newly refreshed attributes
4

Execute Targeted Apply on Verified Non-Breaking Blocks

Once HCL matches live state, run `terraform plan` to confirm that the diff is zero (`No changes. Your infrastructure matches the configuration`). If minor unmanaged resources exist, adopt them cleanly using `terraform import`.

Pro Tip: Golden Rule: Never run blind 'terraform apply' when drift is detected. Use 'terraform apply -refresh-only', inspect HCL differences, and apply code updates until 'No changes' is achieved.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Reconcile drift by backing up state, using 'terraform apply -refresh-only' to pull cloud reality into state without modifies, updating HCL to match reality, and locking resources with lifecycle prevent_destroy."
⚡ 60-Second Elevator Pitch Talking Points
  • Back up remote state immediately with terraform state pull.
  • Inspect the plan to see which arguments are triggering in-place updates vs destructive replacement (-/+).
  • Use lifecycle prevent_destroy and ignore_changes on critical blocks to eliminate downtime risk.
  • Execute terraform apply -refresh-only to sync state with reality, then backport changes into HCL code until plan shows zero diff.
Advertisement
Want more Terraform & IaC scenarios?
Explore our complete collection of scenario-based Terraform & IaC interview runbooks.
Browse All Terraform & IaC Questions →