← Runbook templates

AWS EC2 instance unreachable runbook

Paste into Runspace or any Markdown runbook. {{name}} marks a variable.

Use this runbook when an EC2 instance stops answering SSH, RDP or application traffic. It works through status checks, security groups and network ACLs, the console output and Session Manager, and then helps you decide between a reboot and a stop/start. Every diagnostic step is read-only. Only the recovery steps change anything.

When to use this

  • Health checks, SSH or RDP to one instance time out or get refused, and the rest of the fleet is fine.
  • A CloudWatch alarm on StatusCheckFailed has fired.
  • A deploy, security group edit or subnet change happened shortly before the instance became unreachable.

If several instances in one Availability Zone fail together, check the AWS Health Dashboard before you go further. A single runbook won’t fix a regional issue.

Before you start

  • AWS CLI v2, configured for the target account. Run aws sts get-caller-identity to confirm which account you’re in.
  • IAM permissions for ec2:Describe* and ec2:GetConsoleOutput, plus ssm:DescribeInstanceInformation and ssm:StartSession.
  • The Session Manager plugin for the AWS CLI, if you plan to open a session.
  • Approval rights, or a named approver, for reboot and stop/start.

Variables

  • {{region}}: AWS Region of the instance, e.g. us-east-1.
  • {{instance_id}}: the unreachable instance, e.g. i-0abc123def4567890.
  • {{subnet_id}}: the instance’s subnet. Captured in step 2.
  • {{sg_ids}}: the security group IDs attached to the instance, space-separated. Captured in step 2.
  • {{port}}: the port that’s failing, e.g. 22, 443.

Steps

1. Read status checks

Purpose: find out whether the fault is AWS infrastructure, the guest OS or the attached EBS volumes.

aws ec2 describe-instance-status \
  --region {{region}} \
  --instance-ids {{instance_id}} \
  --include-all-instances \
  --query 'InstanceStatuses[0].{State:InstanceState.Name,System:SystemStatus.Status,Instance:InstanceStatus.Status,Events:Events}'

Expected: State: running, and System and Instance both ok.

Decision: if System is impaired, the underlying host has a problem, so go to step 7 (stop/start). If Instance is impaired, the problem is in the OS, so continue to step 5. If both are ok, the problem is probably the network path: continue to step 2. Scheduled events listed under Events usually explain the outage on their own.

2. Capture network placement

Purpose: get the subnet, security groups and IP addresses that the next steps use.

aws ec2 describe-instances \
  --region {{region}} \
  --instance-ids {{instance_id}} \
  --query 'Reservations[0].Instances[0].{Subnet:SubnetId,SGs:SecurityGroups[].GroupId,PrivateIp:PrivateIpAddress,PublicIp:PublicIpAddress,RootDevice:RootDeviceType}'

Expected: a subnet ID, one or more security group IDs and a private IP. PublicIp will be null if the instance has no public address.

Decision: record {{subnet_id}} and {{sg_ids}}. If RootDevice is instance-store, you can’t stop/start this instance, so step 7 doesn’t apply.

3. Check security group rules

Purpose: confirm that inbound traffic on {{port}} is allowed from where you’re connecting.

aws ec2 describe-security-groups \
  --region {{region}} \
  --group-ids {{sg_ids}} \
  --query 'SecurityGroups[].{Id:GroupId,Inbound:IpPermissions}'

Expected: a rule that covers {{port}} for your source CIDR, or for a security group your traffic comes from.

Decision: if no rule matches, the cause is a security group change. Hand it to the team that owns the group rather than editing it during the incident. Security groups are stateful, so you don’t need to check outbound rules for return traffic.

4. Check network ACLs and routes

Purpose: rule out stateless subnet-level blocks and a missing route.

aws ec2 describe-network-acls \
  --region {{region}} \
  --filters Name=association.subnet-id,Values={{subnet_id}} \
  --query 'NetworkAcls[0].Entries[].[RuleNumber,Egress,Protocol,PortRange.From,PortRange.To,CidrBlock,RuleAction]' \
  --output table

aws ec2 describe-route-tables \
  --region {{region}} \
  --filters Name=association.subnet-id,Values={{subnet_id}} \
  --query 'RouteTables[0].Routes'

Expected: inbound allow on {{port}} and outbound allow on ephemeral ports (1024-65535), with no lower-numbered deny ahead of them. For public access you also need a 0.0.0.0/0 route to an igw- target.

Decision: NACLs are stateless and evaluated in rule-number order, so a deny with a lower number wins. If the route table query returns nothing, the subnet uses the VPC’s main route table. Look that up with Name=association.main,Values=true together with a vpc-id filter.

5. Read the console output

Purpose: look for kernel panics, fsck prompts, a full disk or a failed mount.

aws ec2 get-console-output \
  --region {{region}} \
  --instance-id {{instance_id}} \
  --latest \
  --output text | tail -n 80

Expected: a normal boot that ends at the login prompt or cloud-init completion.

Decision: --latest works on Nitro instances. On Xen instances, leave the flag off. Errors in /etc/fstab, Kernel panic or No space left on device mean the OS needs repair. A reboot rarely fixes those, so escalate with the log attached.

6. Try Session Manager

Purpose: get a shell without going through the network path you’re debugging.

aws ssm describe-instance-information \
  --region {{region}} \
  --filters Key=InstanceIds,Values={{instance_id}} \
  --query 'InstanceInformationList[0].{Ping:PingStatus,Agent:AgentVersion,LastPing:LastPingDateTime}'

aws ssm start-session --region {{region}} --target {{instance_id}}

Expected: Ping: Online, followed by an interactive shell.

Decision: once you’re in, check systemctl status sshd, df -h, free -m and the host firewall (iptables -S or nft list ruleset). If Ping is ConnectionLost, the agent or the OS is hung. Go to step 7.

7. Recover: reboot, or stop and start

Purpose: restart the guest, or move it to new hardware.

A reboot keeps the instance on the same host and keeps its public IP and any instance store data. Use it when the OS is hung and the system status check is ok. A stop/start moves the instance to a new host, which is the fix for an impaired system check. It also erases instance store volumes and changes the public IPv4 address unless an Elastic IP is attached. To check for an Elastic IP, run aws ec2 describe-addresses --filters Name=instance-id,Values={{instance_id}}.

Requires approval: this restarts a production instance and drops any in-flight work.

aws ec2 reboot-instances --region {{region}} --instance-ids {{instance_id}}

Requires approval: this moves the instance to new hardware, erases instance store data and can change the public IP.

aws ec2 stop-instances --region {{region}} --instance-ids {{instance_id}}
aws ec2 wait instance-stopped --region {{region}} --instance-ids {{instance_id}}
aws ec2 start-instances --region {{region}} --instance-ids {{instance_id}}

Verify

aws ec2 wait instance-status-ok --region {{region}} --instance-ids {{instance_id}}
aws ec2 describe-instance-status --region {{region}} --instance-ids {{instance_id}} \
  --query 'InstanceStatuses[0].[SystemStatus.Status,InstanceStatus.Status]'

Both statuses should read ok. Then confirm the real symptom is gone: connect on {{port}}, check that the load balancer target is healthy, and if you did a stop/start, check that DNS still points at the right IP.

Roll back or escalate

A reboot or a stop/start can’t be undone. The rollback is to restore from the latest AMI or EBS snapshot. Escalate to the instance owner and open an AWS Support case if any of these apply:

  • The system check stays impaired after a stop/start.
  • The console output shows filesystem or kernel errors.
  • The instance store-backed instance can’t be recovered with a reboot.

Attach the output from steps 1, 5 and 6 to the case.

Running this in Runspace

In Runspace, each step runs inline and its output streams underneath it. {{subnet_id}} and {{sg_ids}} can be captured from step 2’s output, so nobody pastes IDs by hand. The two recovery commands in step 7 need approval from a separate reviewer, and the approval is pinned to the exact command and runbook revision. Every run is written to a server-side audit log. Runspace is in private pilots: request a pilot with your work email and we’ll reach out to set one up with your team.

FAQ

Should I reboot or stop/start an unreachable EC2 instance?

Reboot when the system status check is ok and the OS looks hung. A reboot keeps the host, the public IP and instance store data. Stop/start when the system status check is impaired. It moves the instance to new hardware, but it erases instance store volumes and changes the public IPv4 address unless an Elastic IP is attached.

Why can't I reach the instance when both status checks pass?

Usually it's the network path: a security group missing an inbound rule for the port, a network ACL deny with a lower rule number, no outbound allow on ephemeral ports, or a missing internet gateway route. A host firewall or a stopped service inside the OS can also cause it. Session Manager lets you check those.

What if Session Manager shows the instance as ConnectionLost?

The SSM agent can't reach the service. Either the OS is hung, or the instance has no route to the SSM endpoints. Read the console output first. If it doesn't show an OS fault, the next step is an approved reboot or stop/start.

Can I stop and start an instance store-backed instance?

No. Instance store-backed instances can only be rebooted or terminated. If a reboot doesn't recover one, restore from the latest AMI and escalate.