How autoscaling works, with AWS Auto Scaling groups
Autoscaling is a feedback loop of metrics, policies and limits. How AWS Auto Scaling groups use target tracking, warm-up, warm pools and health checks, and what doesn't scale.
On this page 11 sections
- The control loop behind autoscaling
- Auto Scaling groups: minimum, maximum and desired
- Target tracking, step and simple policies
- Warm-up, cooldowns and warm pools
- Default instance warmup
- Cooldowns
- Warm pools
- Health checks and instance refresh
- Scaling the parts that don't autoscale
- Key takeaways
- Frequently asked questions
Autoscaling is a feedback loop: a service watches a metric such as average CPU or requests per server, compares it with a target you set, and adds or removes servers to close the gap, always staying between a minimum and a maximum. On AWS, an Auto Scaling group runs this loop for EC2 instances by keeping a "desired capacity" of healthy instances that scaling policies move up and down. Because every step takes minutes, good autoscaling depends as much on warm-up, health checks and a plan for the parts that don't scale as on the policy itself.
The control loop behind autoscaling
Every autoscaler, whether it manages EC2 instances, Kubernetes pods or database replicas, repeats the same four steps:
- Measure. Collect a metric from the group, such as average CPU utilisation or requests per instance.
- Decide. Compare it with a target or threshold and work out how much capacity to add or remove.
- Act. Launch or terminate instances, within the limits you have set.
- Wait. Give new capacity time to start and take traffic before judging the metric again, so the loop doesn't overreact.
The loop is reactive. It responds to load that has already arrived, which is fine when traffic climbs over twenty minutes and too slow when it jumps in twenty seconds. The time goes in several places:
| Stage | What decides how long it takes |
|---|---|
| Metric reaches CloudWatch | EC2 publishes metrics every five minutes by default, or every minute with detailed monitoring |
| Policy decides | How many data points the alarm needs before it acts |
| Instance launches and boots | Instance type, image size and start-up scripts |
| App starts and passes health checks | App start-up time, cache warming and the health check interval |
| Instance counts as warmed up | The warm-up period you configure |
In an illustrative setup with one-minute metrics, a two-minute boot and a one-minute app start, the first new instance serves traffic roughly four to six minutes after a spike begins. Everything in this guide is about making that loop accurate, and then covering the gap it leaves.
Auto Scaling groups: minimum, maximum and desired
An Auto Scaling group is a set of EC2 instances that AWS manages as one unit. You give it a launch template (which image, instance type, network and start-up script to use) and three numbers:
- Minimum: the group never shrinks below this, even at 3 a.m.
- Maximum: the group never grows beyond this, whatever the policies ask for.
- Desired capacity: how many instances the group should run right now. Policies, schedules and people can change it, but only within the minimum and maximum.
The group then does two jobs. It maintains the desired count by replacing instances that fail health checks, and it spreads instances across the Availability Zones you choose, keeping them balanced as it scales. With a minimum of 2, a maximum of 20 and a desired capacity of 4, a policy that asks for 30 instances gets 20, and one that asks for none gets 2.
Treat the maximum as a safety limit, not a formality. It caps your bill if a bug or a bot flood drives traffic up, and it caps how many database connections the fleet can open. If each instance runs four Gunicorn workers with a pool of five connections each, 20 instances can open 400 connections. Set the maximum with that sum in mind.
Target tracking, step and simple policies
A scaling policy decides how far to move the desired capacity. EC2 Auto Scaling has three dynamic policy types, plus scheduled and predictive scaling for load you can see coming.
| Policy | How it decides | Use it when |
|---|---|---|
| Target tracking | You pick a metric and a target value, such as 50% average CPU. AWS creates and manages the CloudWatch alarms and adds or removes capacity in proportion to the gap. | Almost always. AWS recommends it as the default. |
| Step scaling | You create the alarms and define steps, for example add 10% of the group at one breach size and 30% at a larger one. | You want a sharper response at extreme levels, usually alongside target tracking. |
| Simple scaling | One adjustment per alarm, then a cooldown before any further scaling. | Rarely. AWS advises against it. |
Target tracking works like a thermostat, and a few details decide whether it works well:
- The metric must fall as capacity rises. Average CPU and requests per target work, because doubling the instances roughly halves them. Total requests at the load balancer, latency and raw queue length don't. For a queue, divide the backlog by the number of instances.
- Four metrics are built in: average CPU, average network in, average network out, and Application Load Balancer requests per target. You can also use custom metrics.
- Use fast data. Turn on detailed monitoring for one-minute EC2 metrics, or publish high-resolution custom metrics, which target tracking can evaluate at intervals as short as 10 seconds.
- Scale-out is quick and scale-in is gradual. Target tracking removes capacity more slowly than it adds it, to protect availability.
- Policies can be combined. With several target tracking policies, the group scales out if any of them asks, and scales in only when all of them agree.
For a web app behind a load balancer, requests per instance is often a truer signal than CPU; the guide to Kubernetes HPA makes the same case for pods. This policy aims for about 600 requests per instance per minute. The number is illustrative; take yours from a load test.
aws autoscaling put-scaling-policy \
--auto-scaling-group-name web-asg \
--policy-name requests-per-instance \
--policy-type TargetTrackingScaling \
--target-tracking-configuration file://rps.json
# rps.json
{
"TargetValue": 600,
"PredefinedMetricSpecification": {
"PredefinedMetricType": "ALBRequestCountPerTarget",
"ResourceLabel": "app/web-alb/0123abcd/targetgroup/web-tg/4567efab"
}
}
AWS's own target tracking documentation lists the metrics that don't work and why.
Warm-up, cooldowns and warm pools
New instances give misleading readings for their first few minutes. They may show a CPU spike while they boot and fill caches, or near-zero load before the load balancer sends them traffic. Three settings stop the loop from reacting to that noise, or from waiting longer than it needs to.
Default instance warmup
This group-level setting tells Auto Scaling how long after an instance enters service to wait before counting its metrics. While instances warm up, the group scales out further only if the instances that have finished warming are above target, and scale-in is blocked until the new instances are ready. It isn't turned on by default, and AWS strongly recommends setting it. If you don't know the right value, AWS suggests starting at 300 seconds and tuning from there.
Cooldowns
A cooldown pauses further scaling after an activity, for 300 seconds by default. It applies only to simple scaling policies. Target tracking and step scaling don't wait for it and rely on instance warm-up instead, which is one more reason to prefer them.
Warm pools
If your instances take a long time to become ready, perhaps because they download large files or build caches, a warm pool keeps pre-initialised instances beside the group in a stopped, hibernated or running state. A scale-out takes instances from the pool instead of launching from scratch. With stopped instances you pay only for their EBS volumes and any Elastic IP addresses, which is far cheaper than running spare servers all day. AWS also warns that a warm pool you don't need only adds cost, so measure your boot time first. The guide to scaling for exam day covers when pre-warming pays off.
Health checks and instance refresh
By default, the group relies on EC2 status checks: is the instance running, and are its hardware and system checks passing? That misses a common failure: an instance that is up while the app on it has crashed or hung. Turn on Elastic Load Balancing health checks as well, so an instance that fails the load balancer's HTTP check is replaced. Then set a health check grace period that covers boot and app start. The console defaults to 300 seconds, but the AWS CLI and SDKs default to 0.
Keep the health check endpoint shallow. If it queries the database, a slow database makes every instance look unhealthy at once, and replacing healthy servers in the middle of a database incident only adds load. Check that the app process can serve a request, and monitor its dependencies separately.
To roll out a new image or launch template, use an instance refresh. It replaces instances in batches while keeping a minimum healthy percentage of the group in service (90% unless you change it), can pause at checkpoints, and can stop and roll back when a CloudWatch alarm fires. Deployments and scaling then share the same machinery, so a release during busy hours doesn't quietly shrink your capacity.
Scaling the parts that don't autoscale
An Auto Scaling group can add web servers in minutes. Much of what sits behind them can't keep up, and more web servers add pressure on exactly those parts.
| Component | What scales on its own, and what doesn't | What to do |
|---|---|---|
| RDS database | Storage can grow automatically. The instance class never changes by itself, changing it causes downtime, and RDS doesn't add or remove read replicas for you. | Size the primary for peak load, add replicas before known events, and schedule class changes for quiet hours, using Multi-AZ to shorten the outage. |
| Aurora | Aurora Auto Scaling adds and removes reader replicas, never the writer. Aurora Serverless v2 scales each instance's capacity in place. | Send read traffic to readers and set a minimum capacity that holds your working set in memory. |
| Database connections | Every new app instance opens its own connections. RDS for PostgreSQL sets max_connections to LEAST({DBInstanceClassMemory/9531392}, 5000) by default. | Put a pooler in front of the database and cap the group's maximum. |
| Cache nodes | ElastiCache can add shards or replicas through Application Auto Scaling, but each node's memory is fixed. | Size nodes for peak and watch evictions. |
| Third-party APIs | SMS and OTP gateways, payment gateways and email providers enforce their own rate limits. | Queue and retry, and move logins away from the peak minute. |
| Account quotas | EC2 On-Demand quotas are counted in vCPUs per Region. | Check the quota against your maximum before a big event. |
The database row is the one that bites. Autoscaling multiplies connections: 10 instances holding 20 connections each is 200, and at a maximum of 60 instances it becomes 1,200. Connection pooling, with a pooler such as PgBouncer, lets a large fleet share a few dozen real database connections. Horizontal vs vertical scaling explains why the database usually grows by getting bigger rather than by multiplying.
It also helps to see autoscaling as one layer of a wider design. Our guide to handling 100,000 concurrent users shows how caching, queues and a CDN keep most requests from ever reaching the servers you are scaling.
Key takeaways
- Autoscaling is a measure, decide, act and wait loop that runs between a minimum and a maximum you choose.
- Prefer target tracking on a metric that falls as capacity grows, such as requests per instance, with one-minute or faster data.
- Set the default instance warmup and turn on load balancer health checks; they matter as much as the policy.
- Reactive scaling takes minutes, so use schedules or predictive scaling for spikes you can see coming.
- More web servers mean more database connections, so pool connections and cap the group's maximum.
Frequently asked questions
How does autoscaling work in AWS?
On AWS, an EC2 Auto Scaling group holds a desired number of instances between a minimum and a maximum. CloudWatch collects metrics such as CPU or requests per target, scaling policies compare them with a target, and the group launches or terminates instances to match. It also registers instances with the load balancer and replaces unhealthy ones. Other resources, such as ECS services, DynamoDB tables and Aurora replicas, scale the same way through Application Auto Scaling.
What is autoscaling group in AWS?
An Auto Scaling group is a collection of EC2 instances that AWS treats as one logical unit for scaling and management. You define a launch template plus minimum, maximum and desired capacity. The group launches instances to match the desired number, spreads them across Availability Zones and replaces any that fail health checks. Policies, schedules or a person can change the desired capacity, but never beyond the minimum and maximum.
Is autoscaling free?
EC2 Auto Scaling itself has no extra fee. You pay for what it launches and uses: EC2 instances, EBS volumes, load balancer capacity, and CloudWatch features such as detailed monitoring, custom metrics and alarms. A warm pool of stopped instances still costs you for its volumes and Elastic IP addresses. The bigger cost risk is a maximum set too high, which lets a bug or bot traffic scale your bill along with your servers.
Does RDS have autoscaling?
Partly. RDS storage autoscaling grows allocated storage when free space runs low, and it never shrinks it again. RDS doesn't change the instance class for you and doesn't automatically add or remove read replicas. Aurora goes further: Aurora Auto Scaling adjusts the number of reader replicas with target tracking, and Aurora Serverless v2 scales each writer and reader's capacity in small steps without disrupting open connections.