Predictive and scheduled scaling for exam-day traffic spikes

Reactive autoscaling is minutes too late for a 10 a.m. mock test. How to use scheduled actions, predictive scaling and warm pools, and pre-scale databases and caches.

10 min read
On this page 13 sections
  1. Why reactive scaling is too late
  2. Sizing the pre-warm
  3. Scheduled scaling
  4. Predictive scaling
  5. Warm pools and pre-warmed capacity
  6. Pre-scaling the database, cache and CDN
  7. Database
  8. Cache
  9. CDN and load balancer
  10. Services outside your control
  11. A runbook for exam day
  12. Key takeaways
  13. Frequently asked questions

Predictive scaling adds capacity before a forecast rise in traffic, and scheduled scaling adds it at a time you choose. Both exist because reactive autoscaling starts adding servers only after load arrives, which is minutes too late when thousands of students press "Start test" at 10:00. For exam day, raise capacity well before the start with a scheduled action, let predictive scaling handle daily patterns such as evening live classes, and pre-scale the database, cache, CDN and load balancer, which can't grow in seconds.

Why reactive scaling is too late

A mock test doesn't ramp up; it arrives as a wall. Take an illustrative all-India mock for 20,000 students opening at 10:00. If 80% of them press "Start" within 90 seconds and each start triggers five API calls (session check, instructions, paper, first autosave and a timer sync), that is 80,000 requests in 90 seconds, or about 900 a second, against perhaps 50 a second at 9:50. After that, autosaves every 30 seconds keep about 670 requests a second flowing, and a second wall arrives in the last few minutes, when everyone submits.

Reactive autoscaling needs a minute or more of metrics before it acts, then boots instances, starts the app and passes health checks. In an illustrative setup that adds up to four to six minutes. On Kubernetes, the Cluster Autoscaler FAQ puts the chain from rising load to pods running on new nodes at usually about five minutes. Meanwhile the undersized fleet times out, apps retry, and the retries add load.

Time (illustrative)Reactive only, 4 instances at 9:58Pre-scaled to 24 instances at 9:30
9:58Load starts risingIdle headroom
10:00About 900 requests a second hit 4 instances; queues and timeoutsSame load at about 60% of capacity
10:01Alarm fires; first launches beginNormal
10:04 to 10:06New instances start serving, while retries add loadNormal
10:10Still catching up; support phones ringingNormal

Sizing the pre-warm

Pre-scaling only works if you know how much to add. Work it out from a load test rather than a guess:

InputIllustrative value
Peak request rate at the startAbout 900 requests a second
Capacity of one instance at your latency target, from a load test60 requests a second at a 95th-percentile response time under 300 ms
Instances needed at peak900 ÷ 60 = 15
Headroom for uneven arrivals and losing a zoneAbout 50%, giving 22.5
Scheduled minimum24, split evenly across two zones

Keep the scheduled minimum in place until results are out, not just until the paper ends, because the submission wall and result checks are peaks too.

Scheduled scaling

A scheduled action tells an EC2 Auto Scaling group to set its minimum, maximum or desired capacity at a given time, either once or on a recurring cron schedule. Recurring schedules accept an IANA time zone such as Asia/Kolkata, so you can write the times your students see. After the action runs, your normal scaling policies keep working, within the new limits. AWS notes that an action can run up to two minutes late, so don't schedule it for 9:59.

For an illustrative weekly Sunday mock from 10:00 to 13:00, one action raises the floor at 9:30 and another releases it at 13:30:

aws autoscaling put-scheduled-update-group-action \
  --auto-scaling-group-name web-asg \
  --scheduled-action-name sunday-mock-prewarm \
  --recurrence "30 9 * * 0" --time-zone "Asia/Kolkata" \
  --min-size 24 --max-size 80 --desired-capacity 24

aws autoscaling put-scheduled-update-group-action \
  --auto-scaling-group-name web-asg \
  --scheduled-action-name sunday-mock-release \
  --recurrence "30 13 * * 0" --time-zone "Asia/Kolkata" \
  --min-size 4 --max-size 80

On Kubernetes, KEDA's cron trigger does the same for a Deployment: between start and end it keeps at least desiredReplicas pods, and your other triggers can still scale above that. Without KEDA, a scheduled job that raises the HPA's minReplicas works too.

triggers:
- type: cron
  metadata:
    timezone: Asia/Kolkata
    start: 30 9 * * 0
    end: 30 13 * * 0
    desiredReplicas: "24"

Predictive scaling

AWS predictive scaling learns daily and weekly patterns from a CloudWatch metric. It needs at least 24 hours of data, analyses up to the past 14 days, produces an hourly forecast for the next 48 hours and refreshes it every six hours. It starts in forecast-only mode so you can compare the forecast with reality before letting it act. In forecast-and-scale mode it only adds capacity; removing it is left to your dynamic policies. By default it scales at the start of each hour, and a pre-launch setting (SchedulingBufferTime) starts instances earlier so they are warm when the hour begins. The AWS predictive scaling documentation covers the details.

The catch for education platforms is that forecasts only know the past. A daily 7 p.m. live class is a perfect fit. A one-off all-India mock on a new date, or a results announcement, isn't in the history, so the forecast won't see it coming.

Traffic patternBest tool
Same time every day, such as evening live classesPredictive scaling, or a recurring scheduled action
Same time every week, such as a Sunday test seriesA recurring scheduled action; predictive scaling once it has two weeks of history
One-off events, such as a big mock or results dayA one-time scheduled action, sized from a load test
A push notification or WhatsApp blast you send yourselfRaise capacity before you press send
Genuinely unpredictable spikesTarget tracking, a warm pool and a sensible minimum

When several scaling policies are active, AWS sets the desired capacity to the highest any of them asks for, so predictive and target tracking policies can share a group safely, alongside scheduled actions that raise the minimum. The guide to how autoscaling works covers the reactive policies.

Warm pools and pre-warmed capacity

  • Warm pools keep pre-initialised instances beside an Auto Scaling group, usually stopped, so a scale-out skips most of the boot. Stopped instances cost only their EBS volumes and Elastic IPs.

  • Fast images matter as much. Bake dependencies into the machine or container image instead of installing them at boot.

  • Spare node room on Kubernetes comes from low-priority placeholder pods, which real pods displace instantly while a new node boots in the background; see the Kubernetes HPA guide.

  • Guaranteed capacity is available too. EC2 Capacity Reservations hold instance capacity in a specific zone, including future-dated reservations for an event, which require at least 32 vCPUs and a commitment period.

  • Account quotas can quietly cap a scale-out. EC2 On-Demand quotas are counted in vCPUs per Region, so check yours against the scheduled maximum.

Pre-scaling the database, cache and CDN

Database

  • Resize days before, never on the morning. Changing an RDS instance class causes downtime; Multi-AZ shortens it but doesn't remove it.

  • Add read replicas the day before and point reports, leaderboards and dashboards at them.

  • On Aurora Serverless v2, raise the minimum capacity before the event. AWS recommends a minimum that keeps your working set in memory.

  • Warm the buffer cache by running the paper-loading queries once, or with PostgreSQL's pg_prewarm extension, which loads tables and indexes into memory.

  • Check the connection budget at the scheduled maximum, not at today's size, and keep PgBouncer in front of the database.

  • Move heavy jobs out of the window: backups, reindexing, bulk imports and migrations.

Cache

Build the paper into Redis before 10:00 with a warm-up job at 9:30, instead of letting the first request build it. Give keys staggered expiry times so nothing important expires at the start time, and make sure only one request rebuilds a missing key. Our guide to cache stampedes explains the patterns.

CDN and load balancer

Make the app bundle, fonts and question images cacheable with versioned URLs, and load them through the CDN before the start so edge caches already hold them. Dynamic API calls still reach your servers, so the CDN reduces load but doesn't replace pre-scaling. Managed load balancers scale automatically, but for a sudden, unusual spike AWS lets you reserve a minimum Application Load Balancer capacity (an LCU reservation of at least 100 LCUs), sized from the PeakLCUs metric of a load test.

Services outside your control

OTP gateways and payment providers have their own limits. Ask students to log in the day before and keep sessions valid through the exam, so 20,000 OTPs don't go out in the same minute.

A runbook for exam day

  1. A week before: load test at about 1.5 times the expected peak, including the start and submission walls. Fix the first bottleneck and test again.

  2. Three days before: resize the database if needed, check account quotas, create the scheduled actions or cron triggers, and reserve load balancer capacity if the spike is unusual.

  3. The day before: add read replicas, confirm the scheduled actions exist (aws autoscaling describe-scheduled-actions), freeze deployments and brief the on-call team.

  4. 60 minutes before: capacity reaches the scheduled minimum. Check that every instance is healthy and the database has connection headroom.

  5. 30 minutes before: run the cache warm-up and open the lobby, so logins spread over half an hour.

  6. At the start: watch 95th-percentile latency, error rate, database CPU and connections, pooler wait times and queue depth. If something saturates, raise its limit by hand rather than waiting for automation.

  7. Ten minutes before the end: hold capacity for the submission wall, and let scoring run from a queue.

  8. After the exam: scale in on schedule once submissions and results have settled, then review how the forecast compared with the actual traffic.

The institute-side preparation, from exam windows to student communication, is covered in running an online exam for thousands. If you run mock tests on hosted test series software rather than your own servers, use this runbook as a list of questions for your provider before a big mock.

Key takeaways

  • Reactive autoscaling takes minutes; exam traffic arrives in seconds, so capacity must be in place first.

  • Size the pre-warm from a load test: peak request rate divided by per-instance capacity, plus headroom.

  • Use scheduled actions for known events and predictive scaling for daily or weekly patterns.

  • Pre-scale what doesn't autoscale: database size and replicas, warmed caches, CDN content and load balancer capacity.

  • Run the day from a written runbook, and hold capacity until results are out.

Frequently asked questions

What is predictive scaling?

Predictive scaling is autoscaling that acts on a forecast instead of on current load. It studies historical metrics to find repeating patterns, such as a daily evening peak, and adds capacity shortly before the predicted rise, so new servers are ready when traffic arrives. It suits cyclical traffic and applications that take a long time to start. It can't foresee events that aren't in its history, so pair it with schedules for one-off spikes.

What is predictive scaling in AWS?

In AWS, predictive scaling is a policy type for EC2 Auto Scaling groups. It needs at least 24 hours of metric history, analyses up to 14 days, forecasts capacity hourly for the next 48 hours and updates the forecast every six hours. It begins in forecast-only mode. In forecast-and-scale mode it launches instances ahead of the forecast, optionally earlier through a pre-launch buffer, but leaves scaling in to your dynamic policies.

What is scheduled scaling in AWS?

Scheduled scaling in AWS lets you create actions that set an Auto Scaling group's minimum, maximum or desired capacity at a specific time, either once or on a recurring cron schedule with a time zone such as Asia/Kolkata. It is the right tool for known events, such as a weekly mock test. After an action runs, target tracking and other policies continue to scale the group within the new limits.

Share this article

Looking for something else?

Talk to Us