Q: terraform plan shows unexpected changes — what would you investigate?
Deep investigation runbook for unexpected terraform plan diffs: distinguishing configuration drift, provider upgrade defaults, dynamic attributes, and preventing accidental destruction.
#Terraform #State Drift #Plan Diff #AWS Provider #Forces Replacement
🎙️ Candidate Opening & Architectural Context
"When 'terraform plan' shows unexpected changes — especially destruction or in-place modifications — my absolute first rule is: STOP. Do not run 'terraform apply'. Investigate the diff systematically."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Analyze the Diff & Action Symbols
Read the exact plan output carefully:
~ update in-place: Non-destructive attribute change.-/+ replace (forces replacement): High danger — Terraform must DESTROY the existing resource and recreate a new one! Look for the# ... forces replacementcomment next to the specific attribute that triggered it (e.g. EC2 AMI change, VPC CIDR, or EBS volume type).- destroy: Resource is being deleted because it was removed from code or renamed.
2️⃣
Check for Out-of-Band State Drift
Did someone modify the cloud resource manually in the AWS Console or CLI?
- Run:
terraform plan -refresh-only. - This compares the Terraform state file directly against real-world cloud infrastructure without proposing changes to match the code.
- If
-refresh-onlyshows drift: Someone touched the resource out-of-band. - Check AWS CloudTrail: Filter event history by resource name to identify who modified the resource, when, and via which API call.
3️⃣
Provider Upgrades & Variable Drift
Inspect underlying dependencies:
- Provider Version Changes: Check
.terraform.lock.hcl. Did a provider update (e.g. AWS provider v5.0 → v5.2) change default attributes, deprecate tags, or modify schema defaults? - Variable & tfvars Mismatch: Was a different
.tfvarsfile passed? Did an engineer override variables viaTF_VAR_*environment variables? - Dynamic Inputs: Is code using dynamic functions like
timestamp()oruuid()that produce new diffs on every single run?
4️⃣
Resolve and Prevent Recurrence
How to resolve safely:
- If code was renamed: Use
moved { from = ... to = ... }blocks (Terraform 1.1+) instead of destroying and recreating! - If manual change was intentional: Update Terraform code to match reality, run
terraform apply -refresh-only. - If manual change was rogue: Run
terraform applyto overwrite manual changes and restore code enforcement. - Add
lifecycle { prevent_destroy = true }to critical databases and VPCs.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never apply an unexpected plan. Look for '# forces replacement' flags. Run 'terraform plan -refresh-only' to isolate cloud drift from code changes, and check CloudTrail for out-of-band console edits."
⚡ 60-Second Elevator Pitch Talking Points
- Freeze: Do not apply. Inspect plan diff for '~' (in-place) vs '-/+' (destroy and recreate - forces replacement).
- Isolate drift: Run 'terraform plan -refresh-only' to identify out-of-band manual changes in AWS Console.
- Check CloudTrail: Identify who modified the resource outside Terraform and when.
- Check provider & variables: Check '.terraform.lock.hcl' for provider upgrades and verify correct .tfvars file.
- Use moved blocks: If refactoring, use 'moved' blocks to prevent destroy/recreate cycles.
- Safeguards: Add 'lifecycle { prevent_destroy = true }' on critical production resources.
Advertisement