Q: A terraform apply is failing due to configuration drift, but the underlying infrastructure is live and mission-critical. How do you reconcile state and fix the drift without causing downtime?
Production runbook for safely reconciling configuration drift during a failing terraform apply on mission-critical live infrastructure without risking service interruption.
Want to master this scenario in a live sandbox? KodeKloud's HashiCorp Certified Terraform Associate (003) Interactive Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Run Non-Destructive Speculative Plan & Backup Remote State
Immediately create a backup snapshot of the current remote state file before running any write operations. Run `terraform plan` and analyze every resource flagged with `~` (update in-place), `+` (create), and especially `-` or `- / +` (destroy and recreate).
# Backup remote state
terraform state pull > terraform-state-backup-$(date +%s).json
# Run targeted speculative plan without touching live infrastructure
terraform plan -detailed-exitcode -out=tfplan.binary
terraform show -json tfplan.binary > plan_analysis.json
Isolate Why the Resource Threatens Destruction
If Terraform plans to destroy and recreate a live resource, identify the argument forcing recreation (e.g. changing an EC2 AMI, RDS storage type, or VPC subnet CIDR). Add `lifecycle { create_before_destroy = true }` or `lifecycle { ignore_changes = [ ... ] }` in your HCL code to block destructive replacements.
# Prevent accidental destruction of live databases/clusters
lifecycle {
prevent_destroy = true
ignore_changes = [
tags["LastModified"],
allocated_storage,
security_groups
]
}
Refresh State from Reality and Backport Manual Changes into HCL
Run `terraform apply -refresh-only` (in Terraform 1.1+) to update Terraform's state file with real-world infrastructure attributes without applying configuration changes. Once state matches cloud reality, inspect the diff and update your Terraform HCL code to mirror the live configuration.
# Inspect and approve state refresh safely
terraform apply -refresh-only
# Update HCL code to match newly refreshed attributes
Execute Targeted Apply on Verified Non-Breaking Blocks
Once HCL matches live state, run `terraform plan` to confirm that the diff is zero (`No changes. Your infrastructure matches the configuration`). If minor unmanaged resources exist, adopt them cleanly using `terraform import`.
- Back up remote state immediately with terraform state pull.
- Inspect the plan to see which arguments are triggering in-place updates vs destructive replacement (-/+).
- Use lifecycle prevent_destroy and ignore_changes on critical blocks to eliminate downtime risk.
- Execute terraform apply -refresh-only to sync state with reality, then backport changes into HCL code until plan shows zero diff.