Q: Terraform state is huge (200MB) and plan takes 12 minutes. How do you fix it?
Refactoring a 200MB monolithic Terraform state file into decoupled, isolated micro-states to slash execution duration from 12 minutes to under 30 seconds while eliminating blast radius.
🛠️ Production Runbook & Step-by-Step Resolution
Decompose Monolith by Lifecycle Layers
Separate infrastructure by change velocity: 1) Networking (VPCs, Subnets, Transit Gateways — changes once a year), 2) Data Stores (RDS, Elasticache, S3 — changes monthly), 3) Compute & Kubernetes (changes weekly), 4) Application Workloads (changes daily).
Migrate State Resources with Zero Downtime (terraform state mv)
Do NOT destroy and recreate resources. Use terraform state mv or moved blocks in Terraform 1.1+ to migrate resources into separate backend state files without touching live infrastructure.
# Extracting VPC resources to a separate networking state
terraform state mv -state=monolith.tfstate -state-out=networking.tfstate \
aws_vpc.main aws_vpc.main
terraform state mv -state=monolith.tfstate -state-out=networking.tfstate \
aws_subnet.private aws_subnet.private
Inter-State Communication via terraform_remote_state
Connect the layers cleanly using terraform_remote_state data sources or SSM Parameter Store outputs so Compute workspaces can read VPC subnet IDs without sharing state files.
- Decompose monolithic infrastructure into lifecycle layers (Networking, Database, Compute, Apps).
- Use terraform state mv or moved blocks to migrate resources into isolated state files without recreation.
- Connect layer outputs using terraform_remote_state data sources or AWS SSM parameters.
- Result: terraform plan drops from 12 minutes to under 30 seconds with minimal blast radius.