⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Terraform & IaC Interview Questions Scenario 117 of 117 in Terraform & IaC
Senior DevOps / SRE Terraform Drift Reconciliation & Emergency Hotfixes Operations & Support Loop

Q: Suppose an on-call engineer makes a direct change in the AWS Console during a P1 incident — for example, temporarily allowing a port to restore service. How would you handle that change afterward? What if it needs to remain permanently?

Production runbook for safely managing and reconciling manual AWS Console modifications made by on-call engineers during an emergency P1 outage.

#Terraform #AWS #Drift #P1 Incident #terraform import #Hotfix #Governance
🎙️ Candidate Opening & Architectural Context
"In an active P1 crisis, engineers occasionally make imperative changes in the AWS Console (e.g. allowing an ingress port on a security group) to restore customer service immediately. However, leaving manual changes unmanaged creates dangerous configuration drift: the next routine CI/CD `terraform apply` will silently destroy the manual rule and re-trigger the P1 outage. The post-incident response depends on whether the change was temporary or permanent."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's HashiCorp Certified Terraform Associate (003) Interactive Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Scenario A: The Console Change was a Temporary Emergency Hotfix

If the manual port opening or instance resize was only temporary: 1. Complete the underlying fix in application code or network configuration. 2. Verify the permanent fix is operational. 3. Run `terraform apply` in CI/CD. Because Terraform detects the manual console change as drift, Terraform automatically removes the unauthorized manual rule and restores the declared security group baseline. 4. Verify that service remains green after Terraform purges the manual rule.

P1 Console Hotfix→Permanent App/Network Fix Verified→Run Terraform Apply (Removes Hotfix Drift)→Validate Zero Outage
2

Scenario B: The Console Change is Required Permanently

If the new port rule is required permanently in production: 1. Do NOT run `terraform apply` immediately (it will delete the rule and trigger an outage!). 2. Open the Terraform HCL codebase and add the new port rule to the `aws_security_group` resource. 3. Run `terraform plan`. Because the HCL code now matches the manual cloud console reality, the plan output should show: `No changes. Your infrastructure matches the configuration`.

# Step 2: Add permanent rule into HCL
resource "aws_security_group_rule" "payment_webhook" {
  type              = "ingress"
  from_port         = 8443
  to_port           = 8443
  protocol          = "tcp"
  cidr_blocks       = ["192.168.10.0/24"]
  security_group_id = aws_security_group.app_sg.id
}
Advertisement
3

Adopting Brand New Manual Resources (terraform import)

If the engineer created an entirely new resource in the console (e.g. a new Route53 record or S3 bucket): 1. Write the resource block in HCL. 2. Run `terraform import ` to bring it under Terraform state management without re-creating it.

terraform import aws_security_group_rule.payment_webhook sg-0123456789_ingress_tcp_8443_8443_192.168.10.0/24
4

Incident Post-Mortem Action Item

Document the manual change in the incident post-mortem. Revoke direct AWS Console write permissions once the post-mortem action item is merged via PR.

Pro Tip: Golden Rule: Never run blind 'terraform apply' after an incident where console changes were made. Always update HCL code first until 'terraform plan' confirms zero destructive diff.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"If temporary, let Terraform apply revert the drift once the root cause is resolved. If permanent, backport the rule into Terraform HCL code first, verify 'No changes' with terraform plan, and import newly provisioned resources using 'terraform import'."
⚡ 60-Second Elevator Pitch Talking Points
  • If temporary, resolve the underlying application issue, then run terraform apply to cleanly remove the console drift.
  • If permanent, backport the change into Terraform HCL code immediately to prevent apply from deleting it.
  • Verify with terraform plan until it outputs 'No changes. Infrastructure matches configuration'.
  • Use terraform import if brand new cloud resources were provisioned in the console during the fire drill.
Advertisement
Want more Terraform & IaC scenarios?
Explore our complete collection of scenario-based Terraform & IaC interview runbooks.
Browse All Terraform & IaC Questions →