Linux Monitoring Commands: CPU, Memory, I/O, and Network

A practical guide to Linux server monitoring commands; what each column of top, vmstat, iostat, and ss means, and which number actually signals a problem.

7 min Updated 10 Oct 2026

The server responds but it's slow. SSH opens with a delay, you run top and see the load average is at 14, but the CPU is almost idle. This is where most people look for a CPU-hungry process and can't find one, because the problem is somewhere else. This text is a short reference for someone who knows what they're looking for and wants to know exactly what each output column measures.

What load average counts and what it doesn't

The first number in uptime is the average number of processes that have been in the run queue or in uninterruptible wait state (D state) over the past 1, 5, and 15 minutes. The keyword is "uninterruptible." A process blocked on a disk read is also counted in this number, even if it hasn't taken a single second of CPU.

So a load of 14 on a 4-core server doesn't necessarily mean the CPU is saturated. First, check the number of cores:

nproc
grep -c ^processor /proc/cpuinfo

If the load is higher than the number of cores, now you need to figure out where the queue is coming from. This distinction is the whole job.

The top columns that actually matter

In top, press 1 to display each core separately; if one core is at 100% and the rest are idle, the problem is single-threaded, not a lack of overall capacity. The c key shows the full command path, and the H key expands threads.

Look at three columns: %wa (iowait), %si (softirq), and %st (steal). If %wa is above 20, the CPU is the victim, not the culprit. If %st is above zero and steady, the host hypervisor isn't giving you your share, and no tuning inside the machine will fix it; this is where you need to talk to your provider. On a cloud server this number is usually zero, and if it isn't, follow up quickly.

The RES column is the process's actual physical memory, and VIRT is almost meaningless; Java and Go processes show hundreds of megabytes of VIRT without having consumed anything. This is the most misleading column in top.

Memory: why free is almost always low

Look at the output of free -h. The available column is the right metric, not free. Linux keeps page cache in memory and doesn't declare it free, because keeping memory free is wasteful. A server whose free is 200 megabytes and whose available is 6 gigabytes doesn't have a memory problem.

free -h
cat /proc/meminfo | grep -E 'MemAvailable|Dirty|Writeback'

Take the two values Dirty and Writeback seriously. If Dirty stays persistently above a few hundred megabytes, the kernel can't get data to disk fast enough. The result is many processes in D state and high load with an idle CPU. Here the problem is I/O, and you need to go to iostat.

How to spot the OOM killer

dmesg -T | grep -i -E 'oom|killed process'
journalctl -k --since "1 hour ago" | grep -i oom

A line like Out of memory: Killed process 2841 (mysqld) means the kernel killed the process. If MySQL restarts for no reason and there's nothing in its own log, this is very likely the cause. systemd services with Restart=always come right back up, and all you see is a brief outage.

I/O: where most high loads are born

The right tool is iostat from the sysstat package, not top:

iostat -xz 2 5

The -x flag gives extended statistics, and -z omits idle disks. Three columns are decisive:

  • %util: the percentage of time the device was busy. Above 90% on a mechanical disk means saturation. On NVMe this number rarely reaches 100 because it has multiple parallel queues, and %util no longer carries its old meaning.
  • await: the average milliseconds of wait per request, including queue time. On an SSD, under 1ms is normal; above 20ms means you have a problem.
  • aqu-sz: the average queue length. A large number with a large await means the device is the bottleneck, not the application.

To see which process is generating this I/O, run iotop -oPa. The -o flag shows only active processes, -P shows processes instead of threads, and -a shows cumulative amounts.

Here's the mistake people make: many translate high %util into "the disk is failing" and go looking to replace hardware, while an unindexed query or a mysqldump backup without --single-transaction is eating the entire disk. The tell is that the problem repeats at specific hours of the night. First look at the time pattern, then blame the hardware.

Network: ss has replaced netstat

netstat is either not installed on new distributions or is slow. Use ss:

ss -s
ss -tulpn
ss -tan state established | awk '{print $4}' | cut -d: -f2 | sort | uniq -c | sort -rn | head

The first command gives a summary of sockets. If timewait has reached tens of thousands, connections are being opened and closed rapidly, and you should consider net.ipv4.tcp_tw_reuse=1. If SYN-RECV is high, you may be the target of a SYN flood.

The third command counts the high-traffic destination ports. A port with thousands of ESTABLISHED connections usually means a connection pool with no cap or a script that doesn't close connections. To see actual bandwidth, run iftop -nNP or nload; vnstat -l also works for cumulative interface usage.

A short script to record a moment's state

When the server gets slow, you're not online at that moment. Set up a simple loop that dumps the state to a file every 30 seconds:

while true; do
  echo "=== $(date -Is) ===" >> /var/log/snapshot.log
  uptime >> /var/log/snapshot.log
  free -m | head -2 >> /var/log/snapshot.log
  iostat -x 1 2 | tail -20 >> /var/log/snapshot.log
  sleep 30
done

The cost is that it generates a bit of I/O itself, and on small disks you need to watch the file growth. But when the problem comes back two days later, this file is your only witness. If you'd rather use a ready-made tool instead of a manual script, the guide on diagnosing the cause of high server load covers the step-by-step path.

Which tool to pick for which job

NeedToolNote
Instant overviewtop / htophtop has thread breakdown and process tree
Queue and iowaitvmstat 1The b column is the number of blocked processes
Diskiostat -xz 2await matters more than %util on NVMe
Networkss -s, iftopSet netstat aside
Historysar -u -r -d 1 3Requires sysstat to be enabled

If you enable sysstat, sar can also show yesterday's state. This is the only way to see the past, because top only sees the present. Enabling it is one line:

sed -i 's/^ENABLED=.*/ENABLED="true"/' /etc/default/sysstat
systemctl enable --now sysstat

One note about vmstat: the r column is the number of processes in the run queue, and b is the number of blocked processes. If r is large, the CPU is the bottleneck. If b is large, it's the disk. This one distinction solves half of the diagnosis.

For servers with heavy workloads where you want to get I/O right from the start, a dedicated server with NVMe disks brings the await difference from tens of milliseconds down to under one millisecond, and that same heavy query no longer takes down the whole machine. If you're on Linux hosting and can't install the tools because you don't have root access, ask support to install sysstat and iotop; without these two, troubleshooting is almost guesswork.

The next step is clear: right now, run iostat -xz 2 5 on the server and note down await. If it's above 20ms, before anything else go to iotop -oPa. The baseline number you record today is your only reference for comparison tomorrow when the server gets slow.

Frequently asked questions

What load average is normal?

The right metric is the ratio of load to the number of cores, not the number itself. On a 4-core server, a load under 4 means the queue is managed, and above 8 means something is saturated. But if %wa is high, the high load is coming from the disk, and adding CPU won't help at all.

Why does free almost always show little memory?

Because Linux spends free memory on page cache and doesn't count it in the free column. The available column gives an accurate estimate of usable memory. If available is below 10% of total memory and swap is also filling up, then you have a real problem.

What's the difference between %util and await in iostat?

%util tells you what percentage of time the disk was busy, and await tells you how many milliseconds each request took on average, including queue wait time. On NVMe, which has multiple parallel queues, %util can stay low while await has gone up; so await is the more reliable metric.

How do I find out which process is causing network traffic?

ss -tunp shows the owning process of each socket, and iftop -nNP displays the bandwidth of each connection live. For cumulative usage per interface, vnstat -l is enough. If a port has thousands of ESTABLISHED connections, check that one first.

Was this page helpful?