terraform plan Shows Unexpected Changes — Investigation Steps
Deep investigation runbook for unexpected terraform plan diffs: distinguishing configuration drift, provider upgrade defaults, dynamic attributes, and preventin...
Master 106+ battle-tested scenario-based Terraform & IaC interview questions for Senior DevOps, Cloud, and SRE engineers. Includes incident runbooks, STAR talking points, and CLI commands.
Deep investigation runbook for unexpected terraform plan diffs: distinguishing configuration drift, provider upgrade defaults, dynamic attributes, and preventin...
Exhaustive breakdown of configuration drift in Terraform: what happens during refresh, how Terraform resolves divergence, and how to eliminate console access in...
Production-grade remote backend architecture for team collaboration: S3 state storage, DynamoDB state locking, encryption, state isolation, and Atlantis/CI exec...
Enterprise architectural comparison: Directory-based separation with reusable modules vs Terraform Workspaces, multi-account AWS strategy, and code reuse....
Comprehensive architectural guide for scaling Terraform across enterprise engineering teams: layered modules, remote state locking, multi-account isolation, sec...
Use terraform refresh to sync state with actual infrastructure, or more precisely terraform plan -refresh-only to see what would change....
This is a race condition. The last write wins — whichever apply finishes last overwrites the state file. This can cause state corruption ......
1. If using remote backend with versioning (S3 + versioning enabled) — restore the previous version of the state file from S3....
Use multiple provider configurations with aliases or split into multiple workspaces/modules:...
Several reasons:...
Recommended structure:...
1. Fork the module — fork the GitHub repo, apply your patch....
1. Mark outputs as sensitive:...
terraform taint <resource> marks a resource for destruction and recreation on the next terraform apply. Even if nothing in the config cha......
Changing identifier for RDS forces replacement — Terraform deletes the old and creates a new one. That means downtime and potential data ......
Use terraform import:...
Use terraform state mv:...
Workspaces let you maintain multiple state files for the same configuration. terraform workspace new staging creates a staging workspace ......
State handling:...
In versions.tf:...
Several tools:...
Option 1: Default tags (AWS provider)...
terraform_remote_state lets one Terraform module read outputs from another module's state file....
Use the for_each meta-argument with a module:...
- terraform validate — checks syntax and basic configuration correctness without connecting to any APIs. Fast. No credentials needed. Che......
Downloads provider plugins, sets up the backend, downloads modules. Must run before any other command. Run again after changing providers......
Update the version constraint in required_providers. Run terraform init -upgrade. Commit the updated .terraform.lock.hcl. Test with terra......
Terraform will want to destroy it on the next apply. If you want to keep the resource but stop managing it with Terraform, use terraform ......
Reads existing infrastructure. Doesn't create or manage. Example: data "aws_ami" "amazon_linux" finds the latest Amazon Linux AMI ID. U...
Never put credentials in Terraform files. Use environment variables (AWS_ACCESS_KEY_ID), IAM instance profiles (on EC2/ECS), OIDC for CI/......
count creates N identical resources, accessed by index. for_each creates one resource per map key/set element. for_each is preferred — re......
Use depends_on. Terraform infers dependencies from references automatically. Use explicit depends_on only when the dependency isn't captu......
Terragrunt adds DRY configuration for Terraform. Handles: auto-generating backend config per environment, module dependency ordering (run......
Partially applied. Resources created before the failure exist in the cloud AND in state. Resources that failed may exist in cloud but not......
Terratest (Go-based) — write tests that apply the module, verify outputs and real cloud resources, then destroy. Checkov/tfsec for static......
Lock file records exact provider versions and checksums downloaded. Yes, commit it. This ensures all team members and CI use the same pro......
Separate Terraform workspaces/directories per region. Primary region deployed normally. DR region deployed from same modules with DR-spec......
Shows the output values defined in outputs.tf after an apply. Useful for scripting: $(terraform output -raw vpc_id). Can be used to pass ......
Use data "aws_caller_identity" "current" {} → bucket = "my-app-${data.aws_caller_identity.current.account_id}"....
Use the terraform_remote_state data source (with risks noted above) or better: share resource identifiers via SSM Parameter Store. Team A......
Outputs a DOT-format dependency graph of all resources. Visualize with Graphviz. Useful for debugging unexpected destroy ordering or unde......
Use create_before_destroy lifecycle:...
Use small/cheap instance types in tests. Destroy immediately after tests (Terratest handles this). Run tests only on PR, not on every com......
Formats Terraform files to the canonical style. Run terraform fmt -check in CI to fail if code isn't formatted. Run terraform fmt -recurs......
module.vpc.vpc_id — access module A's output from another resource in the same root. If in a separate root module, use terraform_remote_s......
For listener rule changes: create new rule before deleting old. create_before_destroy. For target group changes: add new TG to ALB, shift......
Lists all resources in the current state file. Useful for finding the exact Terraform address of a resource before doing state mv or stat......
Multiple layers: lifecycle { prevent_destroy = true } on critical resources. Pipeline policy that fails if plan contains destroys. AWS Co......
Stores state in a local file (terraform.tfstate). Default if no backend configured. OK for learning but never for production: no locking,......
Use time_sleep resource from the hashicorp/time provider:...
Public repository of Terraform modules and providers at registry.terraform.io. Maintained by community and HashiCorp. Use verified module......
In terraform.tfvars: subnet_ids = ["subnet-abc", "subnet-def"]. In CLI: -var='subnet_ids=["subnet-abc","subnet-def"...
OPA evaluates Terraform plan JSON against Rego policies. Example: deny any plan that creates a publicly accessible S3 bucket. Used in CI ......
Use tfenv (Terraform version manager, similar to nvm for Node). Commit a .terraform-version file in each project. tfenv use automatically......
Interactive REPL for evaluating Terraform expressions. Test functions: > cidrsubnet("10.0.0.0/16", 8, 1) → 10.0.1.0/24. Debug complex exp......
Use a CI/CD system per account (GitHub Actions with OIDC, separate role per account). Shared modules in a central registry. Terragrunt or......
Terraform can't compute the value before applying because it depends on the cloud API's response (e.g., an auto-generated ID, assigned IP......
Use terraform state mv to move resources into module paths. Use moved blocks (Terraform 1.1+) as code-tracked refactoring. Test each move......
Forces resource replacement when another resource changes. Example: replace EC2 instance whenever the launch template changes:...
Run terraform plan on a schedule in CI. If plan shows unexpected changes (someone edited the console), alert via Slack/PagerDuty. Set up ......
Generates or updates the .terraform.lock.hcl file for specific platforms. Useful for CI if the lock was created on Mac but CI runs on Lin......
Use count:...
Pulumi uses general-purpose languages (Python, TypeScript, Go) for IaC instead of HCL. Benefits: native language loops, conditionals, tes......
Use the aws_iam_policy_document data source:...
Skips the interactive confirmation prompt. Use only in CI pipelines after a human has reviewed the plan. Never run with -auto-approve fro......
Treat module input changes as an interface change. First add the new variable while still supporting the old one, and map both to the sam......
The resource addresses changed, so Terraform thinks the old objects disappeared and new ones must be created. Preserve state by moving ad......
Remove the sensitive file from Git tracking, rotate every exposed secret, and replace the workflow with a safer input method such as CI v......
Split the repo into independent root modules, each with its own backend key and pipeline target. In CI, detect changed paths, map them to......
First recreate the backend bucket and locking table if needed. Restore the latest valid state from S3 versioning or backup; if no backup ......
Define shared locals or variables in the root module and pass them explicitly into child modules. A common pattern is a common_tags map p......
Use a moved block in Terraform 1.1+:...
Data sources read existing infrastructure during planning, before new resources are created. If the object does not already exist, the lo......
Pin Terraform and provider versions, commit .terraform.lock.hcl, and run plans in a standard CI environment for the final source of truth......
Export only the values consumers truly need, such as IDs, ARNs, or endpoints. Keep module outputs small and stable because outputs become......
Start with the plan and identify the resource where deletion blocks. Common causes are S3 buckets that still contain objects, security gr......
Reuse the same root-module structure or shared child modules, and keep only the variable values different per environment through tfvars,......
Do not approve from the summary alone. Check whether the changes are expected from the module diff, look specifically for replacements or......
Make the account context explicit in CI and local workflows. Use assume_role with fixed account IDs, print the current caller identity in......
Run Terraform through CI/CD only, store plans as build artifacts, require pull request review plus manual approval before apply, and keep......
The generated password is stored in Terraform state, so state protection matters as much as secret protection. Also be careful with resou......
No. Terraform can express the desired configuration and detect drift, but it cannot stop out-of-band changes by itself. Pair Terraform wi......
Introduce the new output alongside the old one first, keep both during a transition period, and update consumers incrementally. Once all ......
Enforce it in the delivery workflow, not just by convention. Require pull requests to reference a ticket, include the ticket ID in commit......
Split when parts of the infrastructure have different lifecycles, owners, blast radius, or deployment frequency. Examples: shared network......
Backend changes affect where Terraform reads and writes state, so treat them carefully. If only the backend settings changed and state is......
First identify the slow resources and data sources from provider logs or CI timing. Replace broad data-source lookups with explicit input......
Replace broad ignore_changes with a narrow list of specific attributes that are intentionally managed outside Terraform. Run a refresh-on......
Use stable, non-display keys for for_each, such as logical IDs that do not change when labels change. Keep the human-readable name as an ......
Confirm that no Terraform process is still running and that the previous apply is not active in the backend. Then use terraform force-unl......
Avoid a blind replacement. Create a new node group with the desired configuration, allow nodes to join, drain workloads gradually with re......
Read the provider changelog and upgrade guide, then test the change in a lower environment first. Keep the provider version pinned and co......
Pin module sources to immutable versions such as tags or commit SHAs. Use a release process for shared modules, test the new version in n......
Split infrastructure by lifecycle and ownership so one state file does not contain unrelated resources. Avoid storing large rendered temp......
Give each preview environment an isolated backend key or workspace name derived from the PR number, and use strict naming prefixes to avo......
Terraform only changes resources whose configuration or tracked dependencies changed. If a deployment should react to file content, inclu......
Write the target module configuration first, add one import block per resource address, and import in small batches. After each batch, ru......
Inspect the dependency graph and the cloud-side error to find the resource still in use. Terraform usually infers dependencies from refer......
-target can be useful for a narrow recovery action, such as recreating one broken dependency, but it should not become a normal deploymen......
Update required_providers, run terraform init, and use terraform state replace-provider when Terraform needs the provider address in stat......
The root module likely did not pass the aliased provider into the child module, or the child module did not declare the provider configur......
Add variable validation for simple rules and use preconditions or check blocks for rules that depend on computed values. For organization......
Give the variable a clear object type with sensible defaults, and use dynamic blocks only when the nested block should exist. Normalize i......
Shared infrastructure should live in separate root modules and state files from temporary application environments. The app stack can rea......
Revert the Terraform code to the last known good version and run a new plan to see what Terraform will change back. Apply that reviewed r......
Architectural evaluation of Infrastructure as Code: core problems solved (reproducibility, drift elimination, audit trails) contrasted with dangerous anti-patte...