How to handle 100,000 concurrent users

Load balancing, caching, queues, read replicas, connection pooling, autoscaling and load testing: how to build a platform that survives its busiest minute.

9 min read
On this page 15 sections
  1. First, what does "concurrent" really mean?
  2. The shape of a system that scales
  3. Load balancing
  4. Horizontal scaling and autoscaling
  5. Caching: the cheapest request is the one you never make
  6. Queues: absorb bursts and protect the core
  7. The database is usually the real bottleneck
  8. Fix queries first
  9. Read replicas
  10. Connection pooling
  11. Later, if you need it
  12. Fail gracefully
  13. Load testing: prove it before your students do
  14. Key takeaways
  15. Frequently asked questions

To handle 100,000 concurrent users, build layers that each stop work from reaching the next: a CDN for static files and video, a load balancer in front of stateless app servers, caches for repeated reads, queues for slow jobs, and a database protected by read replicas and connection pooling. Then prove it with load tests at peak, because a lakh of users arriving in the same minute behaves very differently from a lakh spread across a day.

"Can the platform handle a lakh students at once?" It's a fair question before a big launch, a results-day live stream or a free all-India mock test. This guide covers each building block, from load balancing to the load testing that proves it all works.

First, what does "concurrent" really mean?

100,000 concurrent users is not 100,000 requests per second. People read, think, write answers and watch videos, and only occasionally ask the server for something. A rule from queueing theory, Little's Law, links the numbers: the number of requests in progress equals the rate at which they arrive multiplied by how long each one takes.

For example, if each of 100,000 active users makes one request every 10 seconds, that's 10,000 requests per second. If each request takes 100 milliseconds to serve, about 1,000 are in flight at any moment. Video segments, meanwhile, should come from the CDN and never touch your application servers at all.

The catch is synchronised moments: "Start test" at 10:00, "Join live class" at 7 p.m., results at noon. Anyone who has tried to book a Tatkal ticket knows the pattern: normal load all day, then everyone at once. Plan for the peak minute, not the daily average.

The shape of a system that scales

Follow a request from a student's phone inwards:

  1. CDN edge: static files, images, scripts, video segments and cacheable pages are served here, close to the user.

  2. Load balancer: spreads the remaining requests across many application servers.

  3. Stateless application servers: any server can handle any request, so capacity grows by adding servers.

  4. Cache: an in-memory store such as Redis or Memcached answers repeated questions without touching the database.

  5. Queues and workers: slow jobs are handed off to run in the background.

  6. Database: a primary for writes, read replicas for reads, and a connection pooler in front.

Each layer's job is to stop as much work as possible from reaching the layer behind it.

Load balancing

A load balancer is the traffic police at a busy junction. It spreads requests across healthy servers, stops sending traffic to servers that fail health checks, and usually handles TLS so the application servers don't have to. Layer 4 balancers route network connections; layer 7 balancers understand HTTP and can route by path or header.

Where you can, avoid "sticky sessions" that pin each user to one server. They make load uneven and turn a single server failure into a wave of lost sessions; keep session data in a shared store instead. Long-lived connections, such as WebSockets for live chat, need extra care: balance them by connection count and drain them gracefully during deployments.

Horizontal scaling and autoscaling

Vertical scaling, buying a bigger server, is simple but hits a ceiling and leaves a single point of failure. Horizontal scaling, adding more servers, has no hard ceiling, but it needs stateless servers: no uploads saved to local disk, no sessions held in memory, no scheduled jobs that assume only one server exists.

Autoscaling adds and removes servers based on CPU, request rates or queue length. Two warnings:

  • New servers aren't instant. Booting, starting the application and warming caches can take minutes, and a spike may be over by then. For known events, schedule extra capacity in advance.

  • More app servers can mean more trouble downstream. Each new server opens its own database connections, so unlimited autoscaling of the app tier can knock over the database. Set sensible maximums, and protect the database with pooling.

Caching: the cheapest request is the one you never make

Cache at every layer where it's safe:

  • Browser: long cache lifetimes for versioned static files.

  • CDN: static assets, video, and public pages such as the course catalogue.

  • Application cache: computed results such as leaderboards, course details and configuration.

Watch out for the cache stampede. When a popular cache entry expires, thousands of requests can hit the database at the same instant to rebuild it. Let one request refresh the value while the others briefly receive the old copy, and add a little randomness to expiry times so entries don't all expire together. And never store personalised data in a shared cache under a key another user could hit.

Queues: absorb bursts and protect the core

Anything the user doesn't need in the immediate response belongs in a queue: emails, SMS and WhatsApp messages, certificates, reports, video processing and test scoring. The request finishes quickly while background workers drain the queue at a steady pace. During a spike the queue grows, instead of your servers falling over.

Make workers safe to retry, set failed messages aside for inspection rather than retrying them forever, and alert on queue length and the age of the oldest message. A queue that grows all day is an outage in slow motion.

The database is usually the real bottleneck

Fix queries first

A query that feels instant in testing can take seconds once tables hold millions of rows and thousands of users run it at once. Add the right indexes, read query plans, paginate long lists, and remove "N+1" patterns, where a page runs one query per item instead of one query for all of them.

Read replicas

Most traffic is reads: course pages, dashboards, reports. Read replicas are copies of the database that serve those reads, leaving the primary free for writes. The trade-off is replication lag, since a replica may be a moment behind. If a student submits an answer and immediately reloads, a lagging replica can show stale data, so route a user's reads of their own recent writes to the primary.

Connection pooling

Database connections are expensive. In PostgreSQL, for instance, every connection is a separate server process with its own memory. If 200 application servers each open 20 connections, the database faces 4,000 of them and will struggle long before it runs out of CPU. A connection pooler such as PgBouncer, or a managed database proxy from your cloud provider, lets thousands of application connections share a small, fixed number of real database connections.

Later, if you need it

Partitioning large tables, sharding data across databases, or moving specific workloads such as chat or analytics to specialised stores all add complexity. Reach for them only when the simpler measures run out.

Fail gracefully

  • Put a timeout on every call to another service, so one slow dependency can't tie up all your servers.

  • Use circuit breakers to stop calling a failing service for a while.

  • Rate-limit per user and per IP address to blunt runaway clients and abuse.

  • Keep feature switches ready so you can turn off non-essential features at peak.

  • For extreme events, a waiting room that admits users in batches beats a site that fails for everyone.

Load testing: prove it before your students do

Tools such as k6, Locust, JMeter and Gatling can simulate thousands of users. What matters is how you use them:

  1. Script real journeys, such as log in, open the dashboard, start a test, save answers and submit, with realistic pauses between steps.

  2. Test against production-sized data. An empty database flatters everything.

  3. Ramp up gradually, find the first bottleneck, fix it and repeat. There is always a next one.

  4. Run a spike test, from zero to peak in about a minute, to mimic a test start.

  5. Run a soak test for several hours to expose memory leaks and slowly exhausted connections.

  6. Watch 95th- and 99th-percentile response times, error rates, database CPU and connections, cache hit rates and queue lengths, not just averages.

  7. Test through the real CDN and load balancer, and warn your providers in advance, since a large test can look like an attack.

Key takeaways

  • Concurrent users are not requests per second. Model your real traffic and plan for synchronised peaks.

  • Push work outwards: CDN first, then caches, then stateless app servers, with queues for anything slow.

  • Scale app servers horizontally, and pre-scale for known events rather than relying on autoscaling alone.

  • Protect the database with good queries, read replicas and connection pooling.

  • Load test real user journeys at peak before the real peak arrives.

Handling 100,000 concurrent users isn't about one heroic server or one magic tool. It's a set of layers, each shielding the next, plus the discipline to measure and rehearse. Build those habits early, and the day a lakh of students log in together becomes a busy day rather than a bad one.

Frequently asked questions

What is the difference between concurrent users and requests per second?

Concurrent users are the people active at the same moment; requests per second is how often their apps ask the server for something. Users spend most of their time reading, thinking or watching video, so the two numbers differ widely. As a worked example, 100,000 active users each making one request every 10 seconds produce about 10,000 requests per second.

How many servers do you need for 100,000 concurrent users?

There's no fixed number. It depends on how many requests those users make, how long each request takes, how much the CDN and caches absorb, and how efficient your database queries are. Model the traffic with Little's Law, then load test real user journeys at peak, adding capacity until response times and error rates stay healthy with margin to spare.

What usually breaks first under heavy load?

Usually the database. Slow queries that were fine on small tables, too many connections from a growing fleet of app servers, and cache stampedes after a popular entry expires are the common culprits. Timeouts, connection pooling, read replicas and careful caching protect it, while queues keep slow work such as notifications and scoring out of the request path.

How do you load test for 100,000 users?

Script realistic journeys, such as logging in, starting a test, saving answers and submitting, with pauses between steps, and run them against production-sized data through the real CDN and load balancer. Ramp up gradually to find each bottleneck, run a spike test from zero to peak in about a minute, and a soak test over several hours.

Share this article

Looking for something else?

Talk to Us