← Runbook templates

Disk full on a Linux host runbook

Paste into Runspace or any Markdown runbook. {{name}} marks a variable.

This runbook frees space on a Linux filesystem that is full or close to it. It shows you what filled the disk, including deleted files that a running process still holds open. It then walks through safe cleanup of the systemd journal and old logs, and ends with checks and prevention. Use it when a disk alert fires or writes fail with No space left on device.

When to use this

  • A disk-usage alert fired for a mount on a Linux host.
  • Applications log ENOSPC or No space left on device.
  • df and du disagree about how much space is used.

Don’t use this runbook for container image storage (/var/lib/docker, /var/lib/containerd), database data directories, or LVM/volume resizing. Those have their own procedures. See “Roll back or escalate”.

Before you start

  • Access: SSH to the host and sudo rights.
  • Tools: df, du, find, sort, numfmt (coreutils/findutils), lsof, journalctl and systemctl (systemd hosts).
  • Context: Find out which service owns the host, and whether anything on it is already in maintenance. If the root filesystem is at 100%, sudo and shell history may fail to write. Keep commands short and don’t create files on that mount.

Variables

  • {{host}}: hostname you are working on.
  • {{mount}}: mount point that is full, e.g. / or /var.
  • {{min_size}}: size threshold for the large-file search, in find syntax, e.g. 500M.
  • {{log_dir}}: log directory to clean, e.g. /var/log/myapp.
  • {{days}}: delete rotated logs older than this many days, e.g. 14.
  • {{journal_max}}: target journal size, e.g. 500M.
  • {{service}}: systemd unit holding a deleted file open, e.g. myapp.service.
  • {{pid}}, {{fd}}: process ID and file descriptor number taken from the lsof output in step 4.

Steps

1. Confirm which filesystem is full and how

Purpose: check whether you are out of blocks or out of inodes.

hostname; df -hT {{mount}}; df -i {{mount}}

Expected: Use% near 100% in the first table, or IUse% near 100% in the second.

Decision: if blocks are full, continue. If inodes are full, step 3 matters most. Look for directories holding millions of small files (session dirs, mail queues, caches).

2. Find the largest directories

Purpose: see where space is going, staying on this one filesystem.

sudo du -xh --max-depth=2 {{mount}} 2>/dev/null | sort -rh | head -25

Expected: a ranked list. /var/log, /var/lib, /tmp and application data directories usually lead.

Decision: write down the top two or three paths. If their total is far below what df reports as used, deleted-but-open files are likely the cause (step 4).

3. Find the largest files

Purpose: name the specific files behind the usage.

sudo find {{mount}} -xdev -type f -size +{{min_size}} -printf '%s\t%p\n' 2>/dev/null \
  | sort -rn | head -20 | numfmt --to=iec --field=1

Expected: sizes and paths, such as an unrotated app.log, core dumps, or old archives.

Decision: separate logs (handled below) from data you don’t own. Don’t delete data files without the owning team.

4. Find deleted files still held open

Purpose: find space that rm did not release because a process still has the file open.

sudo lsof -nP +L1 {{mount}} 2>/dev/null

Expected: rows with NLINK 0 and paths ending in (deleted). The PID and FD columns identify the holder (FD such as 4w means descriptor 4).

Decision: large entries here are reclaimed in step 7. If there are none, skip step 7.

5. Check journal and log usage

Purpose: size the safest cleanup targets.

journalctl --disk-usage
sudo du -sh /var/log/* 2>/dev/null | sort -rh | head -15

Decision: if the journal is several GB, run step 6. If one application log directory dominates, run step 8 against it.

6. Vacuum the systemd journal

Purpose: trim archived journal files down to a target size.

Requires approval: permanently deletes archived journal entries on a production host.

sudo journalctl --vacuum-size={{journal_max}}
journalctl --disk-usage

Expected: Vacuuming done, freed X and a lower usage figure. Vacuum only removes archived files, never the active one.

7. Release space held by deleted-but-open files

Purpose: give back space held by a process from step 4.

Requires approval: truncating a live descriptor or restarting a service affects a running workload.

# Preferred: restart the holder so it reopens its files
sudo systemctl restart {{service}}

# If a restart is not acceptable: truncate the deleted file through /proc
sudo truncate -s 0 /proc/{{pid}}/fd/{{fd}}

Expected: the entry drops out of lsof +L1, or shows size 0. Truncate only files that are deleted logs. Never truncate a deleted database or data file.

8. Remove old rotated logs

Purpose: clear compressed, rotated logs past retention. Do a dry run first.

sudo find {{log_dir}} -xdev -type f \( -name '*.gz' -o -name '*.[0-9]' \) -mtime +{{days}} -print

Requires approval: deletes log files that may be needed for audits or investigations.

sudo find {{log_dir}} -xdev -type f \( -name '*.gz' -o -name '*.[0-9]' \) -mtime +{{days}} -print -delete

Decision: if a live log is the problem, truncate it with sudo truncate -s 0 {{log_dir}}/<file> rather than rm. Removing a log the process holds open just moves the problem to step 4.

9. Prevent recurrence

Purpose: cap journal growth and confirm that logrotate covers the offending logs.

Requires approval: changes host configuration and restarts journald.

sudo mkdir -p /etc/systemd/journald.conf.d
printf '[Journal]\nSystemMaxUse=%s\n' '{{journal_max}}' | sudo tee /etc/systemd/journald.conf.d/size.conf
sudo systemctl restart systemd-journald
sudo logrotate -d /etc/logrotate.conf 2>&1 | grep -A3 '{{log_dir}}'

Expected: the logrotate -d debug output lists rules for {{log_dir}}. If none appear, add a logrotate rule through your configuration management, not by hand.

Verify

df -hT {{mount}}; df -i {{mount}}
sudo lsof -nP +L1 {{mount}} 2>/dev/null | wc -l
systemctl is-active {{service}}

Usage should be below your alert threshold with real headroom, there should be no large deleted-but-open files, and the service should report active. Check the application’s error rate for ENOSPC in the minutes after cleanup.

Roll back or escalate

Deleted logs and vacuumed journal entries can’t be restored, so get approval before steps 6 and 8. If a restart in step 7 leaves {{service}} unhealthy, check journalctl -u {{service}} -n 100 and follow that service’s own runbook.

Escalate to the owning team if:

  • usage comes back within hours (a runaway writer)
  • the space is in a database, container storage or other application data
  • the volume needs to be resized

Record the paths and sizes you found so the root cause can be fixed.

In Runspace, each of these steps runs inline with its output streamed underneath. The steps marked “Requires approval” go to a separate reviewer, and the approval is pinned to the exact command and runbook revision. Values from earlier output, such as {{pid}} and {{fd}} from step 4, can be captured as variables for later steps.

FAQ

Why do df and du report different usage?

Usually because a process still has a deleted file open. du only counts files it can reach by path, but df counts every allocated block. Run lsof +L1 on the mount to find the holder, then restart it or truncate the descriptor through /proc.

Is it safe to rm a large log file to free space?

Not if a process is still writing to it. The space stays allocated until the process closes the file. Truncate it with truncate -s 0 instead, or remove it and then restart the writer.

How do I stop the systemd journal from filling the disk?

Set SystemMaxUse in a drop-in under /etc/systemd/journald.conf.d/ and restart systemd-journald. To shrink it immediately, run journalctl --vacuum-size.

What if the disk is out of inodes rather than space?

df -i will show IUse% near 100%. Look for directories holding very large numbers of small files, such as caches, session stores or mail queues, and clean them with the owning team.