Good server monitoring means you hear about problems from your dashboard before you hear about them from your customers. Bad monitoring means either silence until the site is down, or so many alerts that everyone learns to ignore them. This guide covers what to measure on a typical Linux server, which conditions deserve a 3 a.m. phone call, and which belong in a weekly review instead.
Two kinds of monitoring
It helps to separate two questions:
- Is the service working for users? Checked from outside: can the website be reached, does the login page load, does the API respond in time? This is black-box or synthetic monitoring.
- Why is it (or is it about to stop) working? Measured from inside: CPU, memory, disk, processes, database health, logs. This is white-box monitoring.
You need both. Internal metrics can all look green while a DNS or certificate problem stops anyone reaching you, and an external check alone tells you something is wrong but not what.
Server monitoring essentials: what to watch on every server
CPU
Track overall utilisation and, on virtual machines, steal time (shown as st in top), which indicates the hypervisor is giving your CPU time to other guests. High CPU on its own is not a problem; sustained high CPU combined with slow responses is.
Load average
Load average counts processes running or waiting for CPU and, on Linux, those waiting on disk I/O. Compare it with the number of CPU cores (nproc): a load of 4 on a 4-core server is busy but fine; a sustained 12 is not.
Memory and swap
Linux uses spare memory for disk caching, so "free" memory is often low by design. Watch the available column in free -h instead, and watch for swap activity (si/so in vmstat 1). Constant swapping means the server needs more memory or a process is leaking it. Also watch the kernel log for the out-of-memory (OOM) killer terminating processes.
Disk space and inodes
A full disk stops databases, breaks logging and can corrupt data. Monitor every filesystem, plus inode usage (df -i), which can run out on servers with millions of small files even when space remains.
Disk I/O
High %iowait in top or high utilisation in iostat -x 1 points to storage as the bottleneck, common with busy databases on slow volumes.
Network
Track throughput, errors and connection counts. A sudden jump in connections can mean a traffic spike, a bot, or an attack.
Services and processes
Check that the web server, database, PHP-FPM or application processes, cron and any queue workers are running and responding, not merely present.
Application and dependency metrics
- Web server: requests per second, response times, and the rate of 5xx errors.
- Database: connections in use versus the maximum, slow queries, replication lag, and lock waits.
- Queues and background jobs: queue length and age of the oldest job.
- Certificates: days until SSL/TLS expiry. A quick manual check is available with our SSL checker.
- Backups: time since the last successful backup.
Which alerts matter
The golden rule: an alert that wakes someone must require human action now. Everything else is a report or a ticket. A reasonable starting set:
| Condition | Severity | Suggested starting threshold |
|---|---|---|
| Website or API unreachable from outside | Page immediately | Fails from 2+ locations for 2-3 minutes |
| Error rate (5xx) elevated | Page | Well above normal baseline for 5 minutes |
| Disk will fill soon | Page if within hours; ticket if days | Based on growth rate, not only a fixed percentage |
| Critical service down | Page | Not running or failing health check |
| Certificate expiring | Ticket | Under 14 days (earlier for manual renewals) |
| Backup missed | Ticket, same day | No success in 26 hours for daily jobs |
| High CPU or memory | Warning / review | Sustained above 85-90% for 15+ minutes |
| Replication lag | Page if failover depends on it | Above your RPO |
Thresholds are starting points. Tune them to each server's normal behaviour after a few weeks of data.
Avoiding alert fatigue
- Alert on symptoms, investigate causes. Users feel slow pages and errors, not CPU percentages. Page on symptoms; use resource metrics to diagnose.
- Require duration. A two-second CPU spike is noise. Add a "for 5 minutes" condition.
- Route by severity. Pages to phone, warnings to chat or email, trends to a weekly report.
- Review every page. If an alert fired and nobody needed to act, change or remove it.
- Write a short runbook link into each alert explaining what to check first.
Tools to get started
You do not need an expensive platform to monitor well:
- Prometheus with node_exporter and Grafana: open-source metrics, dashboards and alerting (via Alertmanager). Widely used and flexible.
- Zabbix or Netdata: open-source options with lots of built-in templates.
- Cloud-native tools such as Amazon CloudWatch, Azure Monitor and Google Cloud Monitoring for servers on those platforms.
- Hosted external uptime checkers that test your site from several regions.
An example Prometheus alert rule for low disk space:
- alert: DiskSpaceLow
expr: node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes < 0.10
for: 10m
labels:
severity: warning
annotations:
summary: "Less than 10% disk free on {{ $labels.instance }} {{ $labels.mountpoint }}"
Whatever tools you choose, keep at least one external check independent of the infrastructure it watches; a monitoring server inside the same failed data centre cannot tell you it failed. For round-the-clock monitoring with someone responding to the alerts, see our server management service.
Key takeaways
- Combine external checks (is it working?) with internal metrics (why not?).
- Watch CPU, load, available memory, swap, disk space and inodes, I/O, network and key services.
- Page only for conditions that need immediate human action; everything else is a ticket or report.
- Tune thresholds to real baselines and review every alert that fired.