Q: Someone manually changes Terraform-managed infrastructure — what happens?
Exhaustive breakdown of configuration drift in Terraform: what happens during refresh, how Terraform resolves divergence, and how to eliminate console access in production.
#Terraform #Drift Detection #CloudTrail #State #IAM
🎙️ Candidate Opening & Architectural Context
"This is known as 'Configuration Drift'. What happens depends on whether the resource was modified, added, or deleted, and when the next Terraform execution runs."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
What Happens During Next Terraform Run
During the next terraform plan or apply, Terraform performs a refresh:
- State Refresh: Terraform queries the cloud provider API (e.g. AWS API) for all resources tracked in
terraform.tfstate. - If a tracked resource was modified manually: Terraform compares the real-world state against the code. It flags drift and proposes an in-place update to revert the resource back to the code definition.
- If a tracked resource was deleted manually: Terraform state detects the missing resource and proposes recreating it from scratch.
- If unmanaged sub-resources were added manually: For example, someone manually added tags or security group rules: Depending on resource arguments, Terraform may delete the extra rules, ignore them, or fail with conflict errors.
2️⃣
Two Remediation Paths
The team must choose between restoring code or adopting the change:
- Path A (Enforce Code / Revert Drift): Run
terraform apply. Terraform will overwrite the manual console changes and restore the infrastructure to the version-controlled state of truth. - Path B (Adopt the Manual Change into Code): If the manual change was an emergency hotfix: update the
.tfcode to match the new configuration, runterraform planto verify 0 proposed changes, and commit the code.
3️⃣
How to Prevent It from Happening Again
Enterprise governance safeguards:
- Remove Write Access in Production: Developers and DevOps engineers should have
ReadOnlyAccessin the AWS Production Console. Only the CI/CD pipeline's IAM role should have write permissions. - Automated Drift Detection: Run a scheduled daily CI/CD pipeline (e.g. at 2 AM) executing
terraform plan -detailed-exitcode. If exit code is 2 (drift detected), send an automated alert to Slack/PagerDuty. - Use SCPs (Service Control Policies): Enforce AWS Organizations SCPs preventing direct edits to mission-critical VPC, IAM, or KMS resources.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"On next run, Terraform detects drift during refresh and proposes reverting the manual change back to match code. Prevent it by removing AWS console write access and running automated daily drift detection in CI."
⚡ 60-Second Elevator Pitch Talking Points
- What happens: On next 'terraform plan', refresh detects divergence between cloud reality and state. Terraform proposes reverting manual edits to match code.
- If resource deleted manually: Terraform proposes recreating it.
- Remediation: Either 'terraform apply' to overwrite manual changes, or update code to match reality and commit.
- Root prevention: Restrict IAM - zero human write access in Production AWS Console; only CI/CD IAM role has write access.
- Detection: Schedule daily automated 'terraform plan -detailed-exitcode' in CI; alert team on exit code 2 (drift).
Advertisement