DevOps & Deployment

Zero Downtime Deployment

We set up releases where new instances start and pass health checks before old ones are drained, so users never see an error page while you ship. Rollback is tested as part of the setup.

What zero downtime deployment means

In a basic deployment, the old version stops and the new version starts, and for a few seconds or minutes the site returns errors. Zero downtime deployment removes that gap. The new version is started alongside the old one, checked for health, and only then does traffic move across. The old version is stopped once it has finished its in-flight requests.

The web server part is well understood. The harder part, and the one most teams get wrong, is the database: for a short time both versions run against the same data, so schema changes must work with both.

Who needs it

  • SaaS products used throughout the working day, where there is no quiet window to deploy.
  • E-commerce stores where a brief error during checkout means a lost order.
  • Platforms serving customers across several time zones.
  • Teams that want to deploy several times a day without scheduling maintenance windows.
  • APIs consumed by mobile apps or partners that retry poorly on errors.

What’s included

  • A deployment strategy suited to your setup: rolling, blue-green or container replacement.
  • Health check endpoints that confirm the application is truly ready, not just running.
  • Load balancer or reverse proxy configuration that only routes to healthy instances.
  • Graceful shutdown so in-flight requests and background jobs complete.
  • A migration approach that keeps the database compatible with old and new versions.
  • Automated rollback on failed health checks, plus a manual rollback path.
  • Tested runbook for releases and rollbacks.

We also make sure deploys are observable. During a release you can see which version is serving traffic, whether health checks are passing and whether error rates change, so a problem is spotted in seconds rather than after customer complaints.

How we set it up

  1. Review how the application starts, stops, handles sessions and runs migrations.
  2. Add or improve health checks and graceful shutdown handling.
  3. Configure the deployment strategy in your pipeline and proxy or orchestrator.
  4. Agree a migration pattern with your developers for future schema changes.
  5. Run releases under load in staging, including a deliberate failed deploy and rollback.

We also write down what the team should do when a release does go wrong: how to recognise it, how to roll back and who to tell. Releases become routine partly because everyone knows the plan if they are not.

Rolling, blue-green and canary

StrategyHow it worksGood fit when
RollingInstances are replaced one or a few at a timeYou already run several instances behind a load balancer
Blue-greenA full new environment is started, then traffic switches overYou want an instant switch and an instant way back
CanaryA small share of traffic goes to the new version firstYou have enough traffic and monitoring to judge a partial rollout

On a single server, a simple form of blue-green works well: start the new version on a different port or container, confirm it is healthy, switch the proxy, then stop the old one.

Handling database migrations safely

The reliable approach is often called expand and contract. Instead of changing the schema in one step, you split it into stages that each remain compatible with the running code:

  1. Expand: add new columns or tables without removing anything the old version uses.
  2. Deploy: release code that writes to both old and new structures, or reads from the new with a fallback.
  3. Migrate data: backfill existing records in the background.
  4. Contract: once no running version uses the old structure, remove it in a later release.

It takes a little more planning per change, but it is what makes releases genuinely uneventful.

What affects timeline and cost

Applications that are already stateless and run as several instances are quick to move to zero downtime deploys. More work is needed where sessions or uploads live on local disk, long-running jobs must not be interrupted, WebSocket connections need draining, or the database migration process has to change. Running extra instances briefly during deploys can add a small amount to hosting usage.

Common mistakes

  • Health checks that return success before the app can actually serve requests.
  • Renaming or dropping a column in the same release that stops using it.
  • Killing old instances immediately instead of letting in-flight requests finish.
  • Never testing rollback until a real release goes wrong.

Related: High Availability Setup, CI/CD Pipeline Setup and Monitoring & Logging. See all DevOps & Deployment services, the DevOps guide, or contact us.

Frequently asked questions

Can we get zero downtime deploys on a single server?

Yes, for many applications. Running the new version alongside the old one and switching the proxy gives near-seamless releases on one machine.

Do we need Kubernetes for this?

No. Kubernetes supports rolling updates well, but the same result is achievable with Docker Compose, a load balancer or a managed container service.

What about database migrations?

They need to stay compatible with both versions during the switch. We help your team adopt an expand-and-contract pattern for schema changes.

What happens if the new version is broken?

If health checks fail, traffic never moves to it. If a problem appears after switching, rollback returns traffic to the previous version quickly.

Will users stay logged in during a deploy?

Yes, as long as sessions are stored in a shared place such as Redis or the database, which we check as part of the setup.

Talk to us about zero downtime deployment

Rolling releases so users never hit an error page while you're shipping.

Let's talk

Have something you need built, hosted or fixed?

Tell us what you are trying to do. If we are not the right people for it, we will say so.