On this page
What high availability means
A single server is a single point of failure. If it crashes, needs a reboot or its provider has an issue, your service goes down with it. High availability removes that single point by running your application on more than one node, with a load balancer sending traffic only to healthy ones and a plan for the database so data stays consistent.
It is a trade-off. More servers mean more moving parts, and poorly designed high availability can cause outages of its own. The goal is enough redundancy for your real needs, not the most elaborate architecture possible.
Who needs it
- SaaS products with customers who rely on the application during working hours.
- E-commerce stores where downtime directly means lost orders.
- Platforms with contractual availability commitments to their own customers.
- Booking or clinic systems that staff and patients use throughout the day.
If your application can tolerate a short outage while a server is restored from backup, good backups and monitoring may be enough. We will tell you if high availability is not worth it yet.
What’s included
- Architecture review to find single points of failure in the current setup.
- Load balancer with health checks, removing unhealthy nodes automatically.
- Multiple application nodes built identically from code.
- Shared or synchronised storage for uploads and sessions, so any node can serve any user.
- Database replication or a managed database with failover, where the data layer supports it.
- Monitoring and alerts for node health and replication status.
- A documented and rehearsed failover test.
How we build it
- Agree what availability you need and what failure scenarios matter.
- Make the application stateless enough to run on several nodes, moving sessions and uploads out of local disk.
- Build the nodes, load balancer and database layer, ideally with infrastructure as code.
- Move traffic gradually and monitor behaviour.
- Deliberately take nodes down to test failover, then document what happened and how to recover.
Tools and platforms
Load balancing can use managed options on AWS, Google Cloud, Azure, DigitalOcean or Hetzner, or self-managed Nginx or HAProxy. Databases can use PostgreSQL or MySQL replication, or managed database services with built-in failover. Nodes are defined with Terraform and configured with Ansible or Docker so they stay identical. Kubernetes is an option where you run many services, but it is not required for high availability. Monitoring uses Prometheus and Grafana or the provider’s own tools.
What affects timeline and cost
The biggest factors are how ready the application is to run on several nodes, whether the database needs replication or can move to a managed service, and how many regions or zones you need. Infrastructure costs rise with each additional node and managed service, billed by your provider based on usage.
Where a managed database or load balancer from your cloud provider fits, it often reduces the operational work compared with running those pieces yourself. We compare both options, including the ongoing effort each one implies, before you commit.
Levels of availability
High availability is not all or nothing. We usually discuss it as a series of steps, each removing a specific risk, so you can choose the level that matches your business:
| Level | What it protects against | Complexity |
|---|---|---|
| Single server with backups and monitoring | Data loss; slow detection of outages | Low |
| Warm standby server | Long outages if the main server fails; switchover is manual or scripted | Moderate |
| Several app nodes behind a load balancer | Failure of any single application server | Moderate |
| Replicated or managed database with failover | Database server failure | Higher |
| Multiple zones or regions | Loss of a whole data centre or region | Highest |
Many businesses get most of the benefit from the first three levels. Moving to multi-region setups is a significant step in cost and operational effort, and we would only recommend it where the business case is clear. Each level builds on the previous one, so you can start modestly and extend later.
Common mistakes
- Adding a second web server while the database remains a single point of failure.
- Storing uploads or sessions on local disk, so users see different results on each node.
- Building failover and never testing it until a real outage.
- Over-engineering a small application into a complex cluster that is harder to run than the original server.
Related: Zero Downtime Deployment, Server Monitoring and Backup Solutions. See Cloud & Server Management, the cloud and server guide, or contact us.
Frequently asked questions
Do we need Kubernetes for high availability?
No. Two or more nodes behind a load balancer with a resilient database can provide high availability without Kubernetes.
Can you guarantee the site will never go down?
No one can honestly promise that. High availability reduces the impact of common failures, and we design and test for the scenarios that matter to you.
Is high availability the same as backups?
No. High availability keeps you online when a server fails; backups let you recover lost or corrupted data. You need both.
Will our application need code changes?
Sometimes, mainly to move sessions and uploads off local disk. We identify these early so your developers can plan them.
How do you test failover?
We deliberately stop nodes and, where appropriate, the primary database in a controlled window, then observe how traffic moves and how long recovery takes. The results go into the runbook.
Talk to us about high availability setup
Load balancing and failover so one server going down doesn't take you offline.