Q: You run terraform apply, but the deployment fails after successfully creating some resources. How would you troubleshoot the failure and safely recover?
Incident runbook to recover state cleanly when a terraform apply crashes or fails mid-run after creating partial cloud resources.
Want to master this scenario in a live sandbox? KodeKloud's HashiCorp Certified Terraform Associate (003) Interactive Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Identify the Exact Failure Reason in the Apply Log
Inspect the failure output. Common causes: - **API Permission Errors**: Missing IAM actions (e.g. `AccessDenied` on KMS key creation). - **Invalid Arguments**: Resource naming collisions or unsupported CIDR blocks. - **Cloud Service Quotas**: Hit max VPCs or Elastic IPs in the region. - **Network Timeouts**: Resource provisioning timed out (e.g. RDS Aurora cluster creation).
# Inspect Terraform state to confirm what actually persisted
terraform state list
Back Up Current Remote State Before Any Action
Never run modifications on an unstable state without a snapshot backup.
terraform state pull > terraform.state.backup.$(date +%s).json
Decision: Roll Forward (Preferred) vs. Roll Back
- **Roll Forward (Recommended)**: Fix the root issue (e.g. grant missing IAM permission or update the failing parameter in `.tf` code). Re-run `terraform apply`. Because Terraform is declarative and idempotent, it detects the already-created resources in state and only creates the remaining resources.
- **Roll Back**: If the deployment must be discarded, target the partially created resources and remove them using `terraform destroy -target=
# Surgical targeted cleanup of partially created resources if rolling back
terraform destroy -target=aws_instance.broken_worker
Reconcile Orphaned Resources with terraform import
If a cloud resource was created by the cloud provider API right before network timeout dropped the connection, the resource exists in AWS but is missing from Terraform state. Import the resource into state before running apply again to prevent 'ResourceAlreadyExists' errors.
terraform import aws_s3_bucket.data_bucket prod-analytics-data-bucket-unique
- Back up remote state immediately using terraform state pull.
- Analyze the apply log to identify whether permissions, quotas, or syntax caused the break.
- Prefer Rolling Forward: fix the HCL configuration and re-run terraform apply (declarative idempotency).
- Use terraform import if resources were created in the cloud but missed in state due to network timeouts.