WebSocket scaling for live classes, chat and polls

How to scale WebSockets: connections per server, load balancing long-lived connections, Redis pub/sub fan-out, reconnect storms, and when SSE or polling is simpler.

8 min read
On this page 8 sections
  1. Why WebSockets are hard to scale
  2. Connections per server
  3. Load balancing long-lived connections
  4. Fan-out with pub/sub
  5. Reconnect storms and backoff
  6. WebSockets vs polling vs SSE
  7. Key takeaways
  8. Frequently asked questions

Scaling WebSockets is mostly about long-lived connections: every student in a live class holds one open connection for the whole session, so capacity is measured in concurrent connections and messages per second, not in requests. The usual design is several socket servers behind a load balancer that balances by connection count, a pub/sub layer such as Redis to fan messages out across servers, heartbeats to spot dead connections, and reconnects with jittered backoff so a deploy doesn't turn into a stampede. For one-way updates, Server-Sent Events or plain polling are often simpler.

Why WebSockets are hard to scale

  • Connections are stateful. Each one lives on a single server, so a chat message for a class must reach every server that holds a student of that class.

  • Load stays uneven. A new server starts with no connections, and existing ones don't move to it, so autoscaling doesn't rebalance anything by itself.

  • Fan-out multiplies work. One message to a class of 5,000 is 5,000 sends.

  • Everything in the path has an idle timeout. Nginx closes a proxied WebSocket after 60 seconds without data by default, an AWS Application Load Balancer after 60 seconds, and a Network Load Balancer after 350 seconds for TCP.

  • Failures are collective. A deploy, a crash or a network blip drops thousands of connections, and they all try to reconnect at the same moment.

Connections per server

There is no fixed number of connections a server can hold. It depends on what each connection costs and how busy it is:

  • File descriptors. Every connection uses one, so raise the process limit. Nginx's worker_connections defaults to 512 and counts connections to backend servers as well as clients, so each proxied WebSocket uses two.

  • Memory. Kernel socket buffers, TLS state and your framework's per-connection objects add up, and vary a lot between languages and libraries.

  • CPU. Encryption, message serialisation and fan-out cost far more than holding idle connections.

  • Ports on proxies. Linux's default local port range is 32768 to 60999, roughly 28,000 ports. That caps simultaneous connections from one proxy address to one backend address and port, so large deployments spread traffic over several backend ports or addresses.

  • The accept queue. A reconnect storm arrives all at once; Linux's somaxconn has defaulted to 4096 since kernel 5.4, and was 128 before.

A worked example shows where the real load is. Take an illustrative live class of 5,000 students with chat and polls:

ItemIllustrative valueWhat it means
Open connections5,000Spread over three servers, with room for one to fail
Chat messages coming in20 a second, with slow mode onEach goes to 5,000 students: 100,000 deliveries a second
Chat sent in batches every 250 ms4 batches a second20,000 frames a second instead of 100,000
Poll votes5,000 in 30 secondsAbout 170 writes a second; push only the running tally, once a second
Reconnect after a deploy5,000 within secondsNeeds jittered retries and a cheap way to re-authenticate

Holding 5,000 quiet connections is easy. Delivering 100,000 messages a second is not, which is why batching, slow mode and sending tallies instead of every vote matter more than raw connection limits. Load test with your own message rates and aim for comfortable headroom at twice the expected audience.

Load balancing long-lived connections

A WebSocket starts as an HTTP request asking to upgrade, so it needs a balancer that understands the upgrade. An AWS Application Load Balancer supports WebSockets natively, and each connection stays with the server that accepted it. With Nginx, the Upgrade and Connection headers must be passed explicitly, and least-connections balancing suits long-lived sockets better than round robin:

upstream ws_backend { least_conn; server 10.0.2.11:8000; server 10.0.2.12:8000; }
map $http_upgrade $connection_upgrade { default upgrade; '' close; }

server {
    location /ws/ {
        proxy_pass http://ws_backend;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_read_timeout 120s;
    }
}

Pure WebSocket connections don't need sticky sessions, because each one stays on its server anyway. Sticky sessions matter when a library can fall back to HTTP long-polling: the Socket.IO documentation, for example, requires them when long-polling is enabled, because every polling request for a session must reach the same server. Its alternatives are to disable long-polling or to balance on the client.

Two operational habits keep the fleet balanced. After scaling out, ask a small, random share of clients to reconnect, so the new servers fill up. Before a deploy, stop sending new connections to a server and let its clients move over gradually. Our comparison of L4 vs L7 load balancers covers timeouts and draining in more detail.

Fan-out with pub/sub

Once students are spread across servers, a message needs a way to reach all of them. The standard pattern is publish and subscribe: every socket server subscribes to a channel for each class its clients belong to, a message is published once to the broker, and each server delivers it to its own connected students.

Redis pub/sub is the common broker because it is fast and simple, but it has one property you must design around. Redis documents its pub/sub delivery as at most once: a subscriber that is disconnected when a message is published never receives it. So keep chat history somewhere durable, such as Redis Streams or the database, give every message an ID, and let reconnecting clients fetch what they missed since their last ID. On Redis Cluster, sharded pub/sub (Redis 7.0 and later) keeps messages within one shard instead of broadcasting them to every node.

Most frameworks package this pattern. Socket.IO's Redis adapter publishes every multi-client message to a Redis channel for the other servers to deliver, and Django Channels' channels_redis package provides Redis-backed channel layers. Socket.IO's documentation also warns that the adapter doesn't sign or encrypt messages, so keep Redis on a private network. Video itself should travel on a separate path; see WebRTC vs HLS for live classes.

Reconnect storms and backoff

When 5,000 clients lose their connections at once, the worst thing they can do is retry immediately and in step. Each retry also re-authenticates and re-subscribes, so the recovering servers face a crowd before they are ready. The fix is exponential backoff with random jitter. AWS's widely cited analysis of backoff and jitter found that jittered backoff sharply cuts the load retries put on a server compared with plain exponential backoff, and that "full jitter", a random wait up to an exponentially growing cap, used the least work of the variants it tested:

function reconnectDelay(attempt) {
  const base = 1000, cap = 30000;                  // milliseconds
  const max = Math.min(cap, base * 2 ** attempt);
  return Math.random() * max;                      // "full jitter"
}

Heartbeats catch the connections that die silently, such as a phone switching from Wi-Fi to mobile data. The WebSocket protocol has ping and pong control frames, and browsers answer a server's pings automatically, but page JavaScript can't send protocol-level pings. So let the server ping every 25 seconds or so, comfortably inside the shortest idle timeout on the path, and have the client send a small application-level ping and reconnect if no reply arrives. The Nginx WebSocket guide suggests server pings as the alternative to longer proxy timeouts.

WebSockets vs polling vs SSE

ApproachDirectionWhat it takes to runGood for
Short pollingClient asks every few secondsPlain HTTP; responses can be cachedClass schedules and "results are out" banners
Long pollingServer holds each request open until there is newsMany open requests; some libraries need sticky routingA fallback where WebSockets are blocked
Server-Sent EventsServer to client onlyPlain HTTP; the browser reconnects by itself and resends the last event IDPoll results, leaderboards and announcements
WebSocketsBoth waysUpgrade handling, connection-aware balancing and a pub/sub layerChat, live quizzes and anything interactive

Server-Sent Events are worth a close look. The browser's EventSource reconnects automatically, honours a retry interval sent by the server, and sends a Last-Event-ID header so the server can resume where it left off. One limit matters: over HTTP/1.1, browsers allow only six SSE connections per domain across all tabs, so serve them over HTTP/2. If students mostly listen and rarely send, SSE or polling gives you most of the benefit with far less machinery. For how real-time traffic fits into a wider design, see how to scale a video learning platform.

Key takeaways

  • Plan WebSocket capacity in connections and messages per second; fan-out, not connection count, is usually the bottleneck.

  • Balance by connection count, pass the upgrade headers, and line up idle timeouts with heartbeats.

  • Use pub/sub to reach students on other servers, and a durable store for anything they must not miss.

  • Reconnect with exponential backoff and full jitter, and drain servers gradually on deploys.

  • Use SSE or polling where the client mostly listens.

Frequently asked questions

How many WebSocket connections can a server handle?

There's no fixed number. The operating system caps it through file descriptors and memory, and a proxy in front adds its own limits. Idle connections are cheap; the real constraint is usually CPU for encryption and delivering messages. A server holding thousands of quiet connections may struggle with a few thousand busy ones. Find your limit by load testing with realistic message rates and fan-out, and watching CPU, memory and delivery latency.

How many WebSocket connections per server?

For planning, pick a per-server target from your load test, then run enough servers that losing one still leaves room, because its clients will reconnect to the others at once. For example, if a test shows one server delivering your class's message rate comfortably at 4,000 connections, a 10,000-student event needs at least four servers. Keep proxy limits in mind too, such as Nginx's worker_connections and the ports available to a proxy.

Is WebRTC faster than WebSockets?

For audio and video, usually yes. WebRTC is built for real-time media: it normally runs over UDP, so a lost packet doesn't hold up the ones behind it, and it adapts to network conditions. WebSockets usually run over TCP, where one lost packet delays everything after it. For chat, polls and quiz answers, WebSockets are simpler and fast enough. Live classes commonly use both: WebRTC or HLS for video, WebSockets for interaction.

Share this article

Looking for something else?

Talk to Us