A product mention on television, a festival sale, a viral post: traffic can multiply in minutes. A fixed set of servers either sits mostly idle the rest of the year or falls over at the worst possible moment. Auto scaling addresses this by adding capacity when demand rises and removing it when demand falls, automatically. This article explains how it works, how to configure it sensibly, and the application changes it depends on.
Vertical vs horizontal scaling
There are two ways to add capacity:
- Vertical scaling (scaling up): give a server more CPU and memory. Simple, but there is a ceiling, and resizing usually requires a restart.
- Horizontal scaling (scaling out): add more servers and spread traffic across them with a load balancer. No hard ceiling and no downtime, but the application must be designed for it.
When people talk about auto scaling in the cloud, they almost always mean horizontal scaling: a group of identical instances (AWS Auto Scaling groups, Azure Virtual Machine Scale Sets, Google Cloud managed instance groups) or container replicas (Kubernetes Horizontal Pod Autoscaler) that grows and shrinks.
How auto scaling works
Every auto scaling setup has the same building blocks:
- A template describing a new instance: machine image, size, network, startup script.
- Group limits: minimum, maximum and desired number of instances.
- A metric to watch, such as average CPU, requests per instance or queue length.
- A policy deciding when and by how much to scale.
- A load balancer and health checks so new instances receive traffic only once healthy, and broken ones are replaced.
Types of scaling policy
| Policy | How it works | Good for |
|---|---|---|
| Target tracking | Keep a metric near a target, e.g. average CPU at 50% | Most web applications; simplest to set up |
| Step scaling | Add or remove a set number of instances at thresholds | Workloads needing fine control |
| Scheduled | Change capacity at set times | Known peaks: business hours, planned sales |
| Predictive | Uses historical patterns to scale ahead of demand | Regular daily or weekly cycles |
Target tracking is the right starting point for most teams. Combine it with scheduled scaling before events you know about, such as a marketing campaign launch.
Choosing the right metric
CPU utilisation is the default, and works well for compute-heavy applications. It can mislead, though. An application waiting on a slow database may have low CPU but long response times. Consider:
- Requests per target from the load balancer, for web front ends.
- Queue depth (messages waiting) for background workers.
- Response time, as a supporting alarm rather than a primary scaling metric, because it can rise for reasons more servers will not fix.
A Kubernetes example
For containerised applications, the Horizontal Pod Autoscaler adjusts the number of pods. This manifest keeps average CPU near 60%, with between 2 and 10 replicas:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
The pods must declare CPU requests for utilisation to be calculated, and the cluster itself needs a node autoscaler (or a managed equivalent) to add machines when pods no longer fit.
Making your application ready to scale out
Auto scaling only works if any instance can serve any request and instances can disappear at any moment. That requires:
- Stateless web servers. Store sessions in Redis, a database or signed cookies, not on local disk or in memory.
- Shared file storage. User uploads go to object storage (S3, Azure Blob, Google Cloud Storage), not the instance's disk.
- Fast, automated start-up. Bake software into the machine image or container rather than installing it at boot. If an instance takes ten minutes to become ready, it arrives after the spike.
- A health check endpoint that reports whether the instance can really serve traffic, for example by checking its database connection.
- Graceful shutdown. Instances should finish in-flight requests when the load balancer drains them.
- Configuration from the environment. Nothing should be hand-edited on an individual server.
The database is usually the real limit
Adding web servers is easy; the database behind them is often the bottleneck. More app instances mean more connections and more queries. Prepare by adding caching (Redis or Memcached for repeated queries, a CDN for static and cacheable pages), read replicas for read-heavy workloads, and connection pooling. Our database management service covers tuning of this kind.
Common mistakes
- Minimum of one instance. A single instance means a single point of failure. Run at least two across different availability zones.
- No maximum. A bug or attack can scale you into a very large bill. Always set a sensible maximum and a budget alert.
- Scaling in too aggressively. Use cooldown periods so the group does not flap up and down.
- Never testing. Run a load test with a tool such as k6, Locust or JMeter to watch scaling happen before real customers trigger it.
- Forgetting dependencies. Third-party APIs, payment gateways and email services have their own rate limits.
Caching often reduces the need for scaling in the first place. Check that your pages send sensible Cache-Control headers with our HTTP header checker. For designing the overall setup, see our cloud solutions page.
Key takeaways
- Auto scaling adds and removes instances automatically based on a metric and policy.
- Start with target tracking, add scheduled scaling for known events, and always set minimums and maximums.
- Applications must be stateless, start quickly and expose a real health check.
- Protect the database with caching and replicas, and load-test before peak season.