⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Terraform & IaC Interview Questions Scenario 113 of 117 in Terraform & IaC
Senior Cloud Engineer (L2) Terraform State Recovery & Triage L2 Cloud Screen

Q: You run terraform apply, but the deployment fails after successfully creating some resources. How would you troubleshoot the failure and safely recover?

Incident runbook to recover state cleanly when a terraform apply crashes or fails mid-run after creating partial cloud resources.

#Terraform #IaC #State Recovery #terraform apply #Rollback #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When `terraform apply` fails mid-deployment, Terraform automatically updates the state file with whatever resources were successfully provisioned prior to the error. However, the infrastructure is left in a partial, broken state. Safe recovery requires identifying the exact failure error, refreshing state, fixing the root cause, and re-applying or surgically destroying partial resources."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's HashiCorp Certified Terraform Associate (003) Interactive Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Identify the Exact Failure Reason in the Apply Log

Inspect the failure output. Common causes: - **API Permission Errors**: Missing IAM actions (e.g. `AccessDenied` on KMS key creation). - **Invalid Arguments**: Resource naming collisions or unsupported CIDR blocks. - **Cloud Service Quotas**: Hit max VPCs or Elastic IPs in the region. - **Network Timeouts**: Resource provisioning timed out (e.g. RDS Aurora cluster creation).

# Inspect Terraform state to confirm what actually persisted
terraform state list
2

Back Up Current Remote State Before Any Action

Never run modifications on an unstable state without a snapshot backup.

terraform state pull > terraform.state.backup.$(date +%s).json
Advertisement
3

Decision: Roll Forward (Preferred) vs. Roll Back

- **Roll Forward (Recommended)**: Fix the root issue (e.g. grant missing IAM permission or update the failing parameter in `.tf` code). Re-run `terraform apply`. Because Terraform is declarative and idempotent, it detects the already-created resources in state and only creates the remaining resources. - **Roll Back**: If the deployment must be discarded, target the partially created resources and remove them using `terraform destroy -target=` to avoid deleting unaffected live infrastructure.

# Surgical targeted cleanup of partially created resources if rolling back
terraform destroy -target=aws_instance.broken_worker
4

Reconcile Orphaned Resources with terraform import

If a cloud resource was created by the cloud provider API right before network timeout dropped the connection, the resource exists in AWS but is missing from Terraform state. Import the resource into state before running apply again to prevent 'ResourceAlreadyExists' errors.

terraform import aws_s3_bucket.data_bucket prod-analytics-data-bucket-unique
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Terraform persists successful resources into state before failing. Backup state immediately, fix the error, and roll forward with another apply. If orphaned cloud resources exist, import them with 'terraform import'."
⚡ 60-Second Elevator Pitch Talking Points
  • Back up remote state immediately using terraform state pull.
  • Analyze the apply log to identify whether permissions, quotas, or syntax caused the break.
  • Prefer Rolling Forward: fix the HCL configuration and re-run terraform apply (declarative idempotency).
  • Use terraform import if resources were created in the cloud but missed in state due to network timeouts.
Advertisement
Want more Terraform & IaC scenarios?
Explore our complete collection of scenario-based Terraform & IaC interview runbooks.
Browse All Terraform & IaC Questions →