Zero-downtime deployment: blue-green, canary and rolling

Deploy during a live class without anyone noticing: rolling, blue-green and canary compared, draining connections and WebSockets, feature flags and exam-day deploy freezes.

8 min read
On this page 9 sections
  1. What causes downtime during deploys
  2. Rolling deployments
  3. Blue-green deployments
  4. Canary releases
  5. Draining connections and WebSockets
  6. A worked example: deploying during a live class
  7. Rollbacks, feature flags and exam-day freezes
  8. Key takeaways
  9. Frequently asked questions

Zero-downtime deployment means releasing a new version while every request, video segment and live-class connection keeps working. The three standard strategies are rolling updates, which replace servers a few at a time; blue-green, which switches traffic between two complete environments; and canary releases, which send a small share of traffic to the new version first. All three rely on the same basics: health checks that tell the truth, graceful draining of old servers, and database changes that old and new code can both live with.

What causes downtime during deploys

Downtime during a deploy often has nothing to do with bugs in the new code. It comes from the handover between old and new:

  • Processes killed mid-request. A server stops before finishing the requests it has already accepted.

  • Traffic sent to a server that is leaving. The load balancer keeps routing to an instance for a few seconds after it starts shutting down.

  • Traffic sent to a server that isn't ready. A new instance passes a shallow health check before its connections, caches and configuration are warm.

  • Incompatible database changes. The new code needs a column that doesn't exist yet, or a migration locks a busy table.

  • Breaking API changes. Your Android and iOS apps can't all update at once, so last month's app version is still calling your API.

  • Every WebSocket dropped at once. Thousands of live-class chat connections reconnect together and hammer the servers that remain.

Rolling deployments

A rolling deployment replaces instances in small batches: start new ones, wait until they are healthy, remove old ones, repeat. It is the default strategy for a Kubernetes Deployment, where maxSurge and maxUnavailable both default to 25% of the desired replicas, and a stalled rollout is marked failed after progressDeadlineSeconds, 600 seconds by default. For a web tier that must never lose capacity, set maxUnavailable to 0:

spec:
  replicas: 8
  strategy:
    type: RollingUpdate
    rollingUpdate: {maxSurge: 2, maxUnavailable: 0}
  template:
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: web
          readinessProbe:
            httpGet: {path: /healthz/ready, port: 8000}
          lifecycle:
            preStop:
              exec: {command: ["sleep", "10"]}

The readiness probe keeps a new pod out of the load balancer until it can really serve. Make it check what matters, such as a database connection, without being so strict that a brief blip empties the whole pool. The short preStop sleep gives load balancers time to stop sending traffic before the application receives its stop signal. Rolling updates are cheap and simple, but old and new versions serve side by side for the whole rollout, and rolling back is just another rolling update, not an instant switch. The Kubernetes documentation covers kubectl rollout undo for that.

Blue-green deployments

In a blue-green deployment you run two complete environments. Blue serves students; you deploy the new version to green, test it with real requests, then switch the load balancer so green receives all traffic. If something goes wrong, you switch back, which is the fastest rollback of any strategy.

The costs: you need double capacity during the switch, and all traffic moves at once, so green meets peak load with cold caches unless you warm them first. Both environments usually share one database, so schema changes must work for both versions; see our guide to zero-downtime database migrations. Switch at the load balancer rather than in DNS where you can, because DNS caching makes the cut-over gradual and rollback slow. Our comparison of load balancers explains the layer-7 routing that makes this easy.

Canary releases

A canary release sends a small share of real traffic, say 1% then 5% then 25%, to the new version while the old version handles the rest. At each step you compare the canary's error rate, latency and key student journeys, such as video start and test submission, against the old version, and roll forward or back. Tools such as Argo Rollouts and Flagger automate the steps in Kubernetes; a load balancer with weighted target groups can do the same on virtual machines.

You can also canary by cohort rather than by percentage: staff accounts first, then one customer that has agreed to early releases, then everyone. Mobile apps work the same way: Google Play's staged rollouts release an update to a percentage of users that you raise over time, and Apple's phased release spreads an update to automatic updaters over seven days, pausable if a problem appears.

StrategyExtra capacityRollbackUsers exposed to a bad releaseBest for
RollingA small surgeAnother rollout, minutesGrowing share during the rolloutRoutine releases of stateless services
Blue-greenDouble, brieflyInstant switch backEveryone, but only until you switch backBig releases where fast rollback matters most
CanaryA littleFast: shift the weight backA small, controlled shareRisky changes, where real traffic is the best test

Draining connections and WebSockets

Draining means stopping new work reaching a server while letting its current work finish. The timeouts involved must line up:

  • On AWS, an Application Load Balancer stops sending new requests to a deregistering target and waits up to 300 seconds by default for in-flight requests, a setting called the deregistration delay.

  • Kubernetes sends the container a termination signal after the preStop hook, then force-kills it when the grace period ends, 30 seconds by default.

  • Gunicorn gives workers a graceful timeout, also 30 seconds by default, to finish requests before they are killed.

Make the grace period longer than the preStop delay plus your slowest normal request, and keep the load balancer's draining time within that window; otherwise the platform kills requests the application was still happily finishing.

WebSockets for live-class chat and polls need a different approach, because a connection lasts the whole class and draining can't wait two hours. Instead, the server closes connections deliberately with status 1001, which the WebSocket standard defines as an endpoint "going away", and the app reconnects to another server after a short random delay. Keep chat state in a shared store such as Redis so a reconnection can land anywhere; our guide to scaling WebSockets covers the fan-out side.

A worked example: deploying during a live class

Suppose 3,000 students are in a 7 p.m. live class, spread across eight chat pods, about 375 connections each. The deployment replaces two pods at a time. Each step closes about 750 connections; if the app waits a random one to five seconds before reconnecting, those arrive at roughly 190 a second, which the remaining pods absorb easily. Restart all eight pods together with no jitter, and 3,000 reconnections hit at the same instant, just as the new pods are still warming up. The video itself is unaffected either way, because it comes from the CDN. The numbers are illustrative; the pattern of small batches plus client-side jitter is what matters.

Rollbacks, feature flags and exam-day freezes

Separate deploying code from releasing features. With feature flags, new code ships switched off, you turn it on for staff, then a small cohort, then everyone, and a kill switch turns it off in seconds without a deploy. Rollbacks then become rare, and when you need one it is a flag flip or a redeploy of the previous image, not an emergency.

Two rules keep rollbacks possible. Database migrations must be backward compatible, using the expand-and-contract pattern, so the previous version still works against the new schema. And API changes must tolerate older app versions for as long as they remain in use.

Finally, the most reliable deploy on exam day is no deploy. A simple freeze policy for a coaching platform:

  1. No production deploys from 24 hours before a major test or result announcement until two hours after it ends.

  2. Emergency fixes only, approved by two engineers, with a rollback plan written down first.

  3. Feature flags for anything risky that must ship before the window, turned on after it.

  4. Error budgets and dashboards reviewed before the freeze lifts; our guide to monitoring and SLOs shows how.

Key takeaways

  • Downtime during deploys usually comes from the handover: killed requests, premature traffic, incompatible schemas and mass reconnections.

  • Rolling updates are the cheap default; blue-green gives instant rollback; canaries limit how many users see a bad release.

  • Align preStop delays, grace periods and load-balancer draining, and close WebSockets in small batches with jittered reconnects.

  • Decouple deploy from release with feature flags, keep schema changes backward compatible, and freeze deploys around exams.

Frequently asked questions

What is blue green deployment?

Blue-green deployment runs two identical production environments. One, say blue, serves all users while the new version is deployed and tested on the other, green. Traffic is then switched from blue to green at the load balancer. If problems appear, traffic is switched back to blue, which is still running the old version, making rollback almost instant.

What is zero downtime deployment?

Zero-downtime deployment is releasing a new version of an application without users seeing errors, dropped connections or maintenance pages. It combines a gradual or switchable rollout strategy, such as rolling, blue-green or canary, with health checks, connection draining, backward-compatible database migrations and the ability to roll back quickly if the new version misbehaves.

What is blue green deployment in Kubernetes?

In Kubernetes, blue-green is usually done with two Deployments, one per version, labelled for example version: blue and version: green, and a Service whose selector points at one of them. You deploy green, test it through a separate Service or preview route, then change the main Service's selector to green. Switching the selector back is the rollback. Tools such as Argo Rollouts can automate these steps.

What is the difference between canary and rolling deployment?

A rolling deployment replaces instances batch by batch until all run the new version; the share of traffic on new code simply follows how many instances have been replaced. A canary deliberately holds a small, controlled share of traffic on the new version, compares its metrics against the old one, and only then proceeds. Canaries are slower but expose far fewer users to a bad release.

Share this article

Looking for something else?

Talk to Us