The website is crawling, SSH takes ages to respond, and monitoring shows a load average far above normal. High server load is one of the most common problems sysadmins face, and the cause can be anything from a runaway PHP process to a failing disk to a bot hammering your search page. Rebooting may bring temporary relief but hides the evidence. This guide gives a methodical way to find the real cause on a Linux server.
First, understand what high server load means
Run uptime:
$ uptime
14:02:11 up 41 days, 3:12, 1 user, load average: 12.48, 9.10, 4.37
The three numbers are averages over the last 1, 5 and 15 minutes. On Linux, load counts processes that are running, waiting for CPU, or waiting in uninterruptible sleep, which is usually disk or network storage I/O. That last part is important: high load does not always mean the CPU is busy.
Compare the load with the number of CPU cores (nproc). A load of 8 on an 8-core server means the CPUs are fully occupied; on a 2-core server it means processes are queuing. In the example above, load is rising (1-minute value highest), so the problem is getting worse right now.
Step 1: Classify the bottleneck
top or the friendlier htop gives a first picture. Focus on the CPU summary line:
%Cpu(s): 18.2 us, 4.1 sy, 0.0 ni, 6.3 id, 70.8 wa, 0.0 hi, 0.6 si, 0.0 st
- High
us(user): application code is using the CPU: PHP, Node.js, Java, a database query. - High
sy(system): the kernel is busy, often from heavy networking or many processes being created. - High
wa(I/O wait): CPUs are idle waiting for disk. The example above is I/O-bound. - High
st(steal): on a VM, the hypervisor is giving your CPU time to other tenants. The fix is with the provider or a different instance type.
vmstat 1 shows the same information over time, plus the r (runnable) and b (blocked) process counts and swap activity (si/so).
Step 2: If the CPU is busy
Sort processes by CPU in top (press P) or run:
ps -eo pid,user,%cpu,%mem,etime,cmd --sort=-%cpu | head -15
Common culprits and next steps:
- Many PHP-FPM or Apache workers: usually a traffic issue. Go to Step 5.
- The database (mysqld, postgres): look for slow or unindexed queries. In MySQL,
SHOW FULL PROCESSLIST;shows what is running now; in PostgreSQL, querypg_stat_activity. Enable the slow query log to catch repeat offenders. - A backup, cron job or malware scan running at a bad time: reschedule it or lower its priority with
niceandionice. - An unknown process with an odd name, especially using all CPU constantly: possible cryptocurrency mining malware. Investigate its binary path with
ls -l /proc/PID/exeand treat the server as compromised if it is not yours.
Step 3: If memory is the problem
free -h
ps -eo pid,user,%mem,rss,cmd --sort=-rss | head -15
journalctl -k | grep -i -E "out of memory|oom"
Look at the available column, not "free"; Linux deliberately uses spare RAM as disk cache. If available memory is low and swap is actively being used, the server is thrashing: constantly moving memory to and from disk, which drives up I/O wait and load. Typical causes are too many web server or PHP-FPM workers for the RAM, a memory leak in an application, or a database buffer pool set larger than the machine can afford. Size worker limits so that maximum workers multiplied by memory per worker fits comfortably in RAM alongside the database.
Step 4: If disk I/O is the problem
iostat -xz 1 # from the sysstat package
sudo iotop -o # processes currently doing I/O
In iostat, high %util and high await (milliseconds per request) on a device confirm a storage bottleneck. Check what is reading or writing:
- Database queries scanning whole tables because of missing indexes.
- Backups, log rotation or large file copies.
- Swapping, which shows up as I/O on the swap device.
- A full disk or a filesystem nearly out of inodes (
df -h,df -i). - A failing physical disk; check
dmesgfor I/O errors and runsmartctl -aon physical servers.
On cloud VMs, volumes often have IOPS or throughput limits; your provider's console will show whether you are hitting them.
Step 5: If it is traffic
Sometimes the server is simply doing what it is asked, too many times. Look at the access log to see who is asking:
# Top client IPs in the last 10,000 requests
tail -n 10000 /var/log/nginx/access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head
# Most requested URLs
tail -n 10000 /var/log/nginx/access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head
# Current connection count by state
ss -s
Patterns to look for: one IP or a small range making thousands of requests (a scraper, bot or attack), heavy hits on expensive URLs such as search, login or xmlrpc.php, or a legitimate spike from a campaign. Responses include rate limiting in the web server, blocking abusive IPs at the firewall, adding page caching, or putting a CDN with bot protection in front. Our HTTP header checker shows whether your pages are cacheable.
Step 6: Fix, then prevent
Short-term relief might mean restarting a stuck service, killing a runaway query, or blocking an abusive IP. Then address the root cause:
- Add missing database indexes and optimise slow queries.
- Tune worker counts to the available memory.
- Add caching at the page, object or CDN level.
- Move heavy jobs out of peak hours.
- Right-size the server if it is genuinely under-provisioned for normal load.
Finally, make sure monitoring records CPU breakdown, memory, I/O and request rates over time, so the next incident starts with history rather than guesswork. If load problems keep recurring, our server management and database management teams can investigate and tune the stack.
Key takeaways
- Load average counts CPU and I/O waits; compare it with the number of cores.
- Use the
topCPU line to classify the problem: user, system, I/O wait or steal. - Find the responsible processes, then check memory, disk and traffic patterns.
- Fix the root cause (queries, worker limits, caching, scheduling) rather than just rebooting.