Blog

Cloud articles

Understanding Cloud SLAs and Uptime Guarantees

What a cloud SLA really promises: uptime percentages in minutes, service credits, exclusions, composite availability and how to read the fine print.

4 min read Cloud

Hosting and cloud providers advertise numbers like 99.9% or 99.99% uptime, and it is easy to read them as a promise that your service will stay online. A cloud SLA (service level agreement) is something narrower: a contractual commitment about one service's availability, with a defined remedy, usually a partial refund, if the provider misses it. Knowing what an SLA does and does not cover helps you compare providers honestly and design systems that meet your own uptime goals.

What a cloud SLA contains

Most provider SLAs, whether from AWS, Azure, Google Cloud or a hosting company, share the same structure:

  • Covered service: each product has its own SLA. The virtual machine SLA is separate from the database SLA, which is separate from the storage SLA.
  • Availability commitment: the target percentage, measured over a billing month.
  • Definition of "unavailable": for example, no external connectivity, or error rates above a threshold, for a minimum period.
  • Conditions: often you must use a particular configuration, such as instances spread across two or more availability zones, to qualify for the higher figure.
  • Service credits: the remedy, typically a percentage of that service's monthly charge, increasing in tiers the further availability falls below the target.
  • Exclusions: events that do not count.
  • Claim process: how and by when you must request credits. Credits are rarely applied automatically.

What uptime percentages mean in practice

The difference between "three nines" and "four nines" looks small on paper and is large in practice:

AvailabilityAllowed downtime per month (approx.)Allowed downtime per year (approx.)
99%7.3 hours3.65 days
99.5%3.65 hours1.83 days
99.9%43.8 minutes8.77 hours
99.95%21.9 minutes4.38 hours
99.99%4.4 minutes52.6 minutes
99.999%26 seconds5.3 minutes

Note that monthly measurement resets every month. A provider at 99.9% could, in principle, be down for about 43 minutes every month and still meet its SLA.

Service credits are not compensation

This is the most important point to understand. If an outage costs your online shop a day of sales, the SLA remedy is usually a credit against the affected service's bill, not your lost revenue. If that service costs a modest amount per month, the credit will be modest too. Credits are also capped, commonly at the monthly fee for that service.

So an SLA is a signal of how confident a provider is and how seriously it engineers for availability. It is not insurance. Your protection against downtime comes from your own architecture and, where justified, business interruption insurance.

Common exclusions

Read the exclusions section carefully. Typical ones include:

  • Scheduled maintenance announced in advance (common with hosting companies; less so with the large clouds).
  • Problems caused by your configuration, code, or exceeding quotas.
  • Outages of the public internet or your own network outside the provider's control.
  • Force majeure events.
  • Free tiers, previews and beta features, which usually carry no SLA at all.
  • Single instances that do not meet the required redundant configuration.

Composite SLAs: your system is only as available as its chain

Your application depends on several services in sequence: load balancer, virtual machines, database, DNS, maybe a third-party payment API. If any one fails, your service fails. The theoretical availability of a chain is the product of its parts:

Web tier 99.95%  x  Database 99.9%
= 0.9995 x 0.999
= 0.9985  ->  about 99.85% combined

Adding more dependencies in series lowers the total. Adding redundancy in parallel raises it: two independent components that each fail rarely are much less likely to fail at the same time. This is why multi-zone deployments carry higher SLA figures than single instances.

How to compare provider SLAs

  1. Compare like with like: the same service type and the same configuration (single instance vs multi-zone).
  2. Check the measurement method: who measures, how often, and what counts as down.
  3. Look at the credit tiers and cap.
  4. Check the claim window: you may need to file within a set number of days after the incident, with evidence.
  5. Review status page history: past transparency about incidents tells you more than the headline number.

The large providers publish SLAs per service; see, for example, the AWS Service Level Agreements index.

Setting your own uptime target

Instead of starting from the provider's number, start from your business. How much downtime per month is acceptable for your shop, booking system or customer portal? Then design to meet that, using the provider SLAs as inputs. Practical steps include:

  • Run at least two instances across availability zones behind a load balancer; our cloud solutions page covers this kind of design.
  • Use a managed database with automatic failover (often a "multi-AZ" or "high availability" option).
  • Host DNS with a provider that is itself highly available.
  • Plan maintenance with zero-downtime deployments.
  • Measure real availability independently with external uptime monitoring, so you have your own evidence for SLA claims.

If you offer your own customers an SLA, make sure it is no stronger than what your underlying architecture and providers can support. A server management partner with round-the-clock monitoring can help shorten the incidents that do occur.

Key takeaways

  • A cloud SLA is a per-service commitment with service credits as the remedy, not compensation for lost business.
  • Translate percentages into minutes: 99.9% still allows about 43 minutes of downtime a month.
  • Read exclusions, configuration requirements and claim deadlines.
  • Your real availability is the product of every dependency; redundancy and monitoring are what raise it.

Need help with this?

Netifi helps businesses around the world with Cloud. Tell us what you are working on.