Q: EC2 is running but SSH isn't working — what would you check?
Systematic OSI and AWS-layer troubleshooting methodology when an EC2 instance shows 'Running' but SSH connection fails, covering network timeouts, connection refused, key issues, and SSM rescue.
#AWS #EC2 #SSH #Security Groups #VPC #Systems Manager
🎙️ Candidate Opening & Architectural Context
"When SSH fails while EC2 shows 'Running', the very first step is to observe the exact error message: 'Connection timed out' means a networking/firewall block, whereas 'Connection refused' or 'Permission denied' means you reached the OS but sshd or auth failed."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Differentiate Network Timeout vs Connection Refused
Run verbose SSH: ssh -vvv -i key.pem user@ip to see where the handshake stalls:
- Connection Timed Out: Packet dropped before reaching EC2 (Security Group, Route Table, NACL, or my public IP changed).
- Connection Refused: Packet reached the instance, but no process is listening on port 22 (sshd stopped or crashed).
- Permission Denied (publickey): Network and sshd are working, but authentication credentials or file permissions failed.
2️⃣
AWS Network & VPC Checks (If Timed Out)
Verify the packet path from internet to instance ENI:
- Security Group Inbound Rules: Is port 22 allowed from my current external IP? (Did office VPN or ISP change my public IP?).
- Public IP & Subnet Route Table: Does the subnet route table route
0.0.0.0/0to an Internet Gateway (IGW)? (If private subnet, you must connect via Bastion host or AWS Client VPN). - Network ACLs (NACL): Verify NACL allows inbound port 22 AND allows ephemeral ports (1024–65535) outbound for the return traffic.
- Instance Status Checks: Check AWS Console:
System Status Check(AWS hardware) andInstance Status Check(OS kernel). If Instance check fails, OS is frozen.
3️⃣
OS, Key Pair & Disk Checks (If Refused or Denied)
Verify host-side configuration and authentication:
- Key Permissions: Local private key must have
chmod 400 key.pem(SSH client rejects overly permissive keys). - Correct Username: Amazon Linux (
ec2-user), Ubuntu (ubuntu), Debian (admin), CentOS (centos), RHEL (ec2-user). - Disk Space 100% Full: If the root EBS disk is 100% full, sshd cannot allocate a PTY session or write to
/var/log/auth.log, causing instant drops. - EC2 System Log / Screenshot: In AWS Console, select Actions → Monitor and troubleshoot → Get system log / Get instance screenshot to see kernel panics or boot errors.
4️⃣
Rescue Strategies Without SSH
How senior SREs regain access when SSH is dead:
- AWS Systems Manager (SSM) Session Manager: Connect via browser/CLI (bypasses port 22 and SSH keys entirely using the SSM Agent and IAM role).
- EC2 Serial Console: Connect directly to the serial port if enabled on Nitro instances.
- EBS Volume Detach Rescue: Stop the EC2 instance, detach the root EBS volume, attach it as a secondary drive to a healthy rescue EC2 instance, mount it, fix
/etc/ssh/sshd_configor~/.ssh/authorized_keys, reattach, and start.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Categorize the error immediately: 'Timed out' is AWS network/SG; 'Connection refused' is sshd/port; 'Permission denied' is key/username. Always have AWS SSM Session Manager enabled as a zero-SSH out-of-band management backdoor."
⚡ 60-Second Elevator Pitch Talking Points
- Check error type: 'Timed out' = network/firewall; 'Connection refused' = sshd dead; 'Permission denied' = key/user error.
- If timed out: Check Security Group IP whitelist, subnet route table (IGW attached), NACLs (ephemeral return ports).
- If refused/hung: Check EC2 Instance Status Check, EC2 console screenshot, and system log for kernel panic or 100% disk.
- If permission denied: Verify key permissions (chmod 400), correct OS username (ec2-user vs ubuntu).
- Rescue path: Use AWS SSM Session Manager (no port 22 needed), or detach EBS root volume to a rescue instance to fix config.
Advertisement