WebSocket scaling for live classes, chat and polls
How to scale WebSockets: connections per server, load balancing long-lived connections, Redis pub/sub fan-out, reconnect storms, and when SSE or polling is simpler.
On this page 8 sections
Scaling WebSockets is mostly about long-lived connections: every student in a live class holds one open connection for the whole session, so capacity is measured in concurrent connections and messages per second, not in requests. The usual design is several socket servers behind a load balancer that balances by connection count, a pub/sub layer such as Redis to fan messages out across servers, heartbeats to spot dead connections, and reconnects with jittered backoff so a deploy doesn't turn into a stampede. For one-way updates, Server-Sent Events or plain polling are often simpler.
Why WebSockets are hard to scale
- Connections are stateful. Each one lives on a single server, so a chat message for a class must reach every server that holds a student of that class.
- Load stays uneven. A new server starts with no connections, and existing ones don't move to it, so autoscaling doesn't rebalance anything by itself.
- Fan-out multiplies work. One message to a class of 5,000 is 5,000 sends.
- Everything in the path has an idle timeout. Nginx closes a proxied WebSocket after 60 seconds without data by default, an AWS Application Load Balancer after 60 seconds, and a Network Load Balancer after 350 seconds for TCP.
- Failures are collective. A deploy, a crash or a network blip drops thousands of connections, and they all try to reconnect at the same moment.
Connections per server
There is no fixed number of connections a server can hold. It depends on what each connection costs and how busy it is:
- File descriptors. Every connection uses one, so raise the process limit. Nginx's
worker_connectionsdefaults to 512 and counts connections to backend servers as well as clients, so each proxied WebSocket uses two. - Memory. Kernel socket buffers, TLS state and your framework's per-connection objects add up, and vary a lot between languages and libraries.
- CPU. Encryption, message serialisation and fan-out cost far more than holding idle connections.
- Ports on proxies. Linux's default local port range is 32768 to 60999, roughly 28,000 ports. That caps simultaneous connections from one proxy address to one backend address and port, so large deployments spread traffic over several backend ports or addresses.
- The accept queue. A reconnect storm arrives all at once; Linux's
somaxconnhas defaulted to 4096 since kernel 5.4, and was 128 before.
A worked example shows where the real load is. Take an illustrative live class of 5,000 students with chat and polls:
| Item | Illustrative value | What it means |
|---|---|---|
| Open connections | 5,000 | Spread over three servers, with room for one to fail |
| Chat messages coming in | 20 a second, with slow mode on | Each goes to 5,000 students: 100,000 deliveries a second |
| Chat sent in batches every 250 ms | 4 batches a second | 20,000 frames a second instead of 100,000 |
| Poll votes | 5,000 in 30 seconds | About 170 writes a second; push only the running tally, once a second |
| Reconnect after a deploy | 5,000 within seconds | Needs jittered retries and a cheap way to re-authenticate |
Holding 5,000 quiet connections is easy. Delivering 100,000 messages a second is not, which is why batching, slow mode and sending tallies instead of every vote matter more than raw connection limits. Load test with your own message rates and aim for comfortable headroom at twice the expected audience.
Load balancing long-lived connections
A WebSocket starts as an HTTP request asking to upgrade, so it needs a balancer that understands the upgrade. An AWS Application Load Balancer supports WebSockets natively, and each connection stays with the server that accepted it. With Nginx, the Upgrade and Connection headers must be passed explicitly, and least-connections balancing suits long-lived sockets better than round robin:
upstream ws_backend { least_conn; server 10.0.2.11:8000; server 10.0.2.12:8000; }
map $http_upgrade $connection_upgrade { default upgrade; '' close; }
server {
location /ws/ {
proxy_pass http://ws_backend;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_read_timeout 120s;
}
}
Pure WebSocket connections don't need sticky sessions, because each one stays on its server anyway. Sticky sessions matter when a library can fall back to HTTP long-polling: the Socket.IO documentation, for example, requires them when long-polling is enabled, because every polling request for a session must reach the same server. Its alternatives are to disable long-polling or to balance on the client.
Two operational habits keep the fleet balanced. After scaling out, ask a small, random share of clients to reconnect, so the new servers fill up. Before a deploy, stop sending new connections to a server and let its clients move over gradually. Our comparison of L4 vs L7 load balancers covers timeouts and draining in more detail.
Fan-out with pub/sub
Once students are spread across servers, a message needs a way to reach all of them. The standard pattern is publish and subscribe: every socket server subscribes to a channel for each class its clients belong to, a message is published once to the broker, and each server delivers it to its own connected students.
Redis pub/sub is the common broker because it is fast and simple, but it has one property you must design around. Redis documents its pub/sub delivery as at most once: a subscriber that is disconnected when a message is published never receives it. So keep chat history somewhere durable, such as Redis Streams or the database, give every message an ID, and let reconnecting clients fetch what they missed since their last ID. On Redis Cluster, sharded pub/sub (Redis 7.0 and later) keeps messages within one shard instead of broadcasting them to every node.
Most frameworks package this pattern. Socket.IO's Redis adapter publishes every multi-client message to a Redis channel for the other servers to deliver, and Django Channels' channels_redis package provides Redis-backed channel layers. Socket.IO's documentation also warns that the adapter doesn't sign or encrypt messages, so keep Redis on a private network. Video itself should travel on a separate path; see WebRTC vs HLS for live classes.
Reconnect storms and backoff
When 5,000 clients lose their connections at once, the worst thing they can do is retry immediately and in step. Each retry also re-authenticates and re-subscribes, so the recovering servers face a crowd before they are ready. The fix is exponential backoff with random jitter. AWS's widely cited analysis of backoff and jitter found that jittered backoff sharply cuts the load retries put on a server compared with plain exponential backoff, and that "full jitter", a random wait up to an exponentially growing cap, used the least work of the variants it tested:
function reconnectDelay(attempt) {
const base = 1000, cap = 30000; // milliseconds
const max = Math.min(cap, base * 2 ** attempt);
return Math.random() * max; // "full jitter"
}
Heartbeats catch the connections that die silently, such as a phone switching from Wi-Fi to mobile data. The WebSocket protocol has ping and pong control frames, and browsers answer a server's pings automatically, but page JavaScript can't send protocol-level pings. So let the server ping every 25 seconds or so, comfortably inside the shortest idle timeout on the path, and have the client send a small application-level ping and reconnect if no reply arrives. The Nginx WebSocket guide suggests server pings as the alternative to longer proxy timeouts.
WebSockets vs polling vs SSE
| Approach | Direction | What it takes to run | Good for |
|---|---|---|---|
| Short polling | Client asks every few seconds | Plain HTTP; responses can be cached | Class schedules and "results are out" banners |
| Long polling | Server holds each request open until there is news | Many open requests; some libraries need sticky routing | A fallback where WebSockets are blocked |
| Server-Sent Events | Server to client only | Plain HTTP; the browser reconnects by itself and resends the last event ID | Poll results, leaderboards and announcements |
| WebSockets | Both ways | Upgrade handling, connection-aware balancing and a pub/sub layer | Chat, live quizzes and anything interactive |
Server-Sent Events are worth a close look. The browser's EventSource reconnects automatically, honours a retry interval sent by the server, and sends a Last-Event-ID header so the server can resume where it left off. One limit matters: over HTTP/1.1, browsers allow only six SSE connections per domain across all tabs, so serve them over HTTP/2. If students mostly listen and rarely send, SSE or polling gives you most of the benefit with far less machinery. For how real-time traffic fits into a wider design, see how to scale a video learning platform.
Key takeaways
- Plan WebSocket capacity in connections and messages per second; fan-out, not connection count, is usually the bottleneck.
- Balance by connection count, pass the upgrade headers, and line up idle timeouts with heartbeats.
- Use pub/sub to reach students on other servers, and a durable store for anything they must not miss.
- Reconnect with exponential backoff and full jitter, and drain servers gradually on deploys.
- Use SSE or polling where the client mostly listens.
Frequently asked questions
How many WebSocket connections can a server handle?
There's no fixed number. The operating system caps it through file descriptors and memory, and a proxy in front adds its own limits. Idle connections are cheap; the real constraint is usually CPU for encryption and delivering messages. A server holding thousands of quiet connections may struggle with a few thousand busy ones. Find your limit by load testing with realistic message rates and fan-out, and watching CPU, memory and delivery latency.
How many WebSocket connections per server?
For planning, pick a per-server target from your load test, then run enough servers that losing one still leaves room, because its clients will reconnect to the others at once. For example, if a test shows one server delivering your class's message rate comfortably at 4,000 connections, a 10,000-student event needs at least four servers. Keep proxy limits in mind too, such as Nginx's worker_connections and the ports available to a proxy.
Is WebRTC faster than WebSockets?
For audio and video, usually yes. WebRTC is built for real-time media: it normally runs over UDP, so a lost packet doesn't hold up the ones behind it, and it adapts to network conditions. WebSockets usually run over TCP, where one lost packet delays everything after it. For chat, polls and quiz answers, WebSockets are simpler and fast enough. Live classes commonly use both: WebRTC or HLS for video, WebSockets for interaction.