Blog

Servers & Hosting articles

Server Monitoring: What to Watch and Which Alerts Matter

A practical guide to server monitoring: the metrics that matter, sensible alert thresholds, avoiding alert fatigue, and tools to get started.

4 min read Servers & Hosting

Good server monitoring means you hear about problems from your dashboard before you hear about them from your customers. Bad monitoring means either silence until the site is down, or so many alerts that everyone learns to ignore them. This guide covers what to measure on a typical Linux server, which conditions deserve a 3 a.m. phone call, and which belong in a weekly review instead.

Two kinds of monitoring

It helps to separate two questions:

  • Is the service working for users? Checked from outside: can the website be reached, does the login page load, does the API respond in time? This is black-box or synthetic monitoring.
  • Why is it (or is it about to stop) working? Measured from inside: CPU, memory, disk, processes, database health, logs. This is white-box monitoring.

You need both. Internal metrics can all look green while a DNS or certificate problem stops anyone reaching you, and an external check alone tells you something is wrong but not what.

Server monitoring essentials: what to watch on every server

CPU

Track overall utilisation and, on virtual machines, steal time (shown as st in top), which indicates the hypervisor is giving your CPU time to other guests. High CPU on its own is not a problem; sustained high CPU combined with slow responses is.

Load average

Load average counts processes running or waiting for CPU and, on Linux, those waiting on disk I/O. Compare it with the number of CPU cores (nproc): a load of 4 on a 4-core server is busy but fine; a sustained 12 is not.

Memory and swap

Linux uses spare memory for disk caching, so "free" memory is often low by design. Watch the available column in free -h instead, and watch for swap activity (si/so in vmstat 1). Constant swapping means the server needs more memory or a process is leaking it. Also watch the kernel log for the out-of-memory (OOM) killer terminating processes.

Disk space and inodes

A full disk stops databases, breaks logging and can corrupt data. Monitor every filesystem, plus inode usage (df -i), which can run out on servers with millions of small files even when space remains.

Disk I/O

High %iowait in top or high utilisation in iostat -x 1 points to storage as the bottleneck, common with busy databases on slow volumes.

Network

Track throughput, errors and connection counts. A sudden jump in connections can mean a traffic spike, a bot, or an attack.

Services and processes

Check that the web server, database, PHP-FPM or application processes, cron and any queue workers are running and responding, not merely present.

Application and dependency metrics

  • Web server: requests per second, response times, and the rate of 5xx errors.
  • Database: connections in use versus the maximum, slow queries, replication lag, and lock waits.
  • Queues and background jobs: queue length and age of the oldest job.
  • Certificates: days until SSL/TLS expiry. A quick manual check is available with our SSL checker.
  • Backups: time since the last successful backup.

Which alerts matter

The golden rule: an alert that wakes someone must require human action now. Everything else is a report or a ticket. A reasonable starting set:

ConditionSeveritySuggested starting threshold
Website or API unreachable from outsidePage immediatelyFails from 2+ locations for 2-3 minutes
Error rate (5xx) elevatedPageWell above normal baseline for 5 minutes
Disk will fill soonPage if within hours; ticket if daysBased on growth rate, not only a fixed percentage
Critical service downPageNot running or failing health check
Certificate expiringTicketUnder 14 days (earlier for manual renewals)
Backup missedTicket, same dayNo success in 26 hours for daily jobs
High CPU or memoryWarning / reviewSustained above 85-90% for 15+ minutes
Replication lagPage if failover depends on itAbove your RPO

Thresholds are starting points. Tune them to each server's normal behaviour after a few weeks of data.

Avoiding alert fatigue

  • Alert on symptoms, investigate causes. Users feel slow pages and errors, not CPU percentages. Page on symptoms; use resource metrics to diagnose.
  • Require duration. A two-second CPU spike is noise. Add a "for 5 minutes" condition.
  • Route by severity. Pages to phone, warnings to chat or email, trends to a weekly report.
  • Review every page. If an alert fired and nobody needed to act, change or remove it.
  • Write a short runbook link into each alert explaining what to check first.

Tools to get started

You do not need an expensive platform to monitor well:

  • Prometheus with node_exporter and Grafana: open-source metrics, dashboards and alerting (via Alertmanager). Widely used and flexible.
  • Zabbix or Netdata: open-source options with lots of built-in templates.
  • Cloud-native tools such as Amazon CloudWatch, Azure Monitor and Google Cloud Monitoring for servers on those platforms.
  • Hosted external uptime checkers that test your site from several regions.

An example Prometheus alert rule for low disk space:

- alert: DiskSpaceLow
  expr: node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
        / node_filesystem_size_bytes < 0.10
  for: 10m
  labels:
    severity: warning
  annotations:
    summary: "Less than 10% disk free on {{ $labels.instance }} {{ $labels.mountpoint }}"

Whatever tools you choose, keep at least one external check independent of the infrastructure it watches; a monitoring server inside the same failed data centre cannot tell you it failed. For round-the-clock monitoring with someone responding to the alerts, see our server management service.

Key takeaways

  • Combine external checks (is it working?) with internal metrics (why not?).
  • Watch CPU, load, available memory, swap, disk space and inodes, I/O, network and key services.
  • Page only for conditions that need immediate human action; everything else is a ticket or report.
  • Tune thresholds to real baselines and review every alert that fired.

Need help with this?

Netifi helps businesses around the world with Servers & Hosting. Tell us what you are working on.