Q: An EC2 instance is running normally, and the application is also working as expected. However, the application team has raised a Sev-1 incident because they are unable to SSH into the EC2 instance. You have joined the Sev-1 call. What areas would you check to restore SSH connectivity?
Sev-1 incident response procedure when an EC2 instance application operates normally, but operations engineers cannot establish SSH sessions.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Check Security Group & Network ACL (NACL) Ingress Rules
Verify port 22 ingress rules on the EC2 instance's Security Group. Check if someone modified the corporate VPN CIDR or accidentally deleted the port 22 rule. Verify that the subnet NACL allows inbound port 22 and outbound ephemeral ports (1024-65535).
# Verify Security Group rules via AWS CLI
aws ec2 describe-security-groups \
--group-ids sg-0123456789abcdef0 \
--query "SecurityGroups[*].IpPermissions[?ToPort==\`22\`]"
Verify Routing Table & Bastion / VPN Path
If engineers connect via a Bastion host or AWS Client VPN, verify that the Bastion instance is operational and its route table directs traffic through the Internet Gateway or Transit Gateway.
aws ec2 describe-route-tables --filters "Name=association.subnet-id,Values=subnet-01234567"
Bypass SSH via AWS Systems Manager (SSM) Session Manager
Do NOT waste precious Sev-1 time debugging port 22 if you need immediate terminal access. Connect directly using AWS SSM Session Manager. SSM relies on the outbound SSM agent and does not require port 22, public IPs, or security group ingress rules.
# Instantly connect via SSM bypassing SSH entirely
aws ssm start-session --target i-0123456789abcdef0
Inspect Host sshd Daemon, iptables, and Disk Space
Once inside via SSM or EC2 Serial Console: 1. Check `systemctl status sshd` (is sshd running or listening on an alternate port?). 2. Check `journalctl -u sshd -n 50` for authentication failures or MaxStartups drops. 3. Check `df -h` (if root `/` is 100% full, sshd cannot allocate pty sessions and immediately terminates connections).
- Check Security Group and Subnet NACL ingress rules for port 22 and corporate VPN CIDRs.
- Use AWS SSM Session Manager to immediately bypass SSH and access the terminal within seconds.
- Use the EC2 Serial Console if the network stack is completely unresponsive.
- Inspect sshd daemon status, root disk space (/tmp and /), and host iptables/firewalld rules.