Q: How would you identify unused or over-provisioned cloud resources?
Actionable audit runbook for discovering and terminating ghost cloud resources: unattached EBS volumes, orphan managed disks, idle load balancers, unassociated Elastic IPs, and obsolete snapshots.
#FinOps #AWS #Azure #Cost Explorer #Orphaned Disks #KubeCost
🎙️ Candidate Opening & Architectural Context
"In every production cloud account over 1 year old, 15 to 25% of the monthly spend consists of 'zombie' or orphaned resources left behind by deleted clusters, test VMs, or failed CI pipelines."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
The Top 5 Orphaned Cloud Culprits
Where ghost money leaks every month:
- Unattached EBS Volumes / Azure Managed Disks: When an EC2/VM is deleted, attached persistent volumes often remain (status
availableor unattached), continuing to bill per GB-month. - Unassociated Elastic IPs / Public IPs: Cloud providers charge an hourly penalty fee for allocated public IPs that are NOT attached to a running instance.
- Idle Load Balancers (ALB/NLB): Abandoned load balancers with 0 healthy targets or 0 request count that bill base hourly fees (~$25+/mo each).
- Old Snapshots & AMIs: EBS snapshots from deleted instances retained for years without retention lifecycles.
- Orphaned NAT Gateways: Idle NAT Gateways left running in obsolete testing VPCs billing ~$35/mo base + data transfer.
2️⃣
Discovery Tools & Automated Auditing
How senior teams automate identification:
- AWS Compute Optimizer & Cost Explorer: Flags over-provisioned EC2/RDS instances and underutilized EBS volumes.
- Azure Advisor: Generates automated Cost recommendations for idle virtual network gateways, unused disks, and downsized VMs.
- KubeCost / OpenCost: In-cluster real-time allocation tool that breaks down Kubernetes spend by namespace, deployment, and orphaned persistent volumes.
- Cloud Custodian: Open-source policy engine running automated custodial crons to tag and auto-terminate unattached volumes older than 7 days.
3️⃣
Quick CLI Audit Commands
Run immediately to find leaks:
- AWS Unattached EBS:
aws ec2 describe-volumes --filters Name=status,Values=available --query 'Volumes[*].[VolumeId,Size,CreateTime]' --output table - AWS Unassociated IPs:
aws ec2 describe-addresses --query 'Addresses[?NetworkInterfaceId==null].[PublicIp,AllocationId]' --output table - Azure Unattached Disks:
az disk list --query '[?managedBy==null].[name,resourceGroup,diskSizeGb]' -o table
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Implement continuous cloud hygiene: Audit for unattached EBS/managed disks, unassociated public IPs, and zero-target load balancers using AWS Compute Optimizer, Azure Advisor, and CLI query filters. Automate cleanup with Cloud Custodian policies."
⚡ 60-Second Elevator Pitch Talking Points
- Top leaks: Unattached EBS/managed disks, unassociated Elastic IPs, idle ALBs with 0 targets, and forgotten NAT Gateways.
- Discovery tools: AWS Cost Explorer / Compute Optimizer, Azure Advisor Cost blade, and KubeCost for in-cluster visibility.
- Fast CLI checks: 'aws ec2 describe-volumes --filters Name=status,Values=available' to instantly find orphan storage.
- Automated prevention: Cloud Custodian or Lambda janitor scripts to notify Slack and terminate unattached disks after 7 days.
- Enforce mandatory tagging ('Environment', 'Owner', 'Project') in Terraform so untagged rogue resources cannot be created.
Advertisement