Horizontal vs vertical scaling: which should you use?
Scale up or scale out? The trade-offs in cost, limits and downtime, what stateless apps need, and why databases usually grow vertically first, with a worked cost example.
On this page 9 sections
Vertical scaling (scaling up) gives one machine more CPU, memory or faster storage; horizontal scaling (scaling out) adds more machines and spreads the work across them. Scale up while it is cheap and simple, and scale out once you need more than one machine can give, protection from a single failure, or the ability to follow a daily peak. Most platforms do both: application servers scale out, while the database scales up for as long as it sensibly can.
The difference at a glance
| Question | Vertical (scale up) | Horizontal (scale out) |
|---|---|---|
| How you add capacity | Move to a bigger machine | Add machines behind a load balancer |
| Ceiling | The largest machine you can rent | Practically none for stateless services |
| Change without downtime? | Usually no; resizing means a restart | Yes; machines join and leave while traffic flows |
| A machine fails | The service is down | The others carry the load |
| Changes to your app | None | The app must not keep state on the server |
| Following a daily peak | Poor; you pay for peak size all day | Good, especially with autoscaling |
| Typical fit | Databases, early-stage apps, stateful services | Web and API servers, background workers, video processing |
Vertical scaling
Vertical scaling is the upgrade everyone understands: the same software on a bigger computer. On AWS, for example, the general-purpose M7i family runs from m7i.large, with 2 vCPUs and 8 GiB of memory, up to m7i.48xlarge, with 192 vCPUs and 768 GiB. Nothing in your code changes, and everything stays in one memory space, which keeps latency low and operations simple.
The limits show up in three places:
- Resizing means downtime. AWS says you must stop an EC2 instance before changing its instance type, and changing an RDS instance class causes downtime too, though a Multi-AZ deployment shortens it.
- One machine is one point of failure. A hardware fault, a kernel panic or a bad deploy takes everything down with it.
- A bigger box only helps software that can use it. A single-threaded process uses one core however many you buy. PgBouncer, for example, is single-threaded, so a larger server needs several PgBouncer processes to use its extra cores.
Horizontal scaling
Horizontal scaling puts several identical servers behind a load balancer. Capacity grows in small steps, a failed server is simply replaced, and servers can be spread across data centres so that one outage doesn't stop the service. With autoscaling, the number of servers can follow traffic through the day, which a single big machine can't do without restarts.
The price is coordination. Requests cross the network more often, deployments must handle a mixed fleet, and anything a server used to keep to itself now has to live somewhere shared.
What horizontal scaling demands of your app
An app that runs happily on one server often breaks in subtle ways on three. Before scaling out, check that:
- Sessions live in a shared store such as Redis or the database, not in a server's memory, so any server can handle any request without sticky sessions.
- Uploads go to object storage, not to a local disk that other servers can't see.
- Scheduled jobs run once, through a scheduler or a lock, instead of once per server. Otherwise every server sends the 7 a.m. reminder SMS.
- Caches are shared or allowed to be briefly stale. Per-server caches will disagree with each other for a while.
- Risky requests are idempotent. A retried test submission or payment callback must not create a duplicate.
- Servers are identical and disposable, built from the same image and configured through environment variables.
- Logs and metrics leave the machine, because the server that failed may no longer exist.
- Database connections are budgeted. Servers multiplied by pool size must fit what the database, or its PgBouncer pooler, can take.
- Real-time messages reach every server. Chat and live polls need a pub/sub layer, because two students in one class may be connected to different servers.
Databases: the special case
Stateless app servers are easy to multiply. A database isn't, because it is the shared source of truth: in a single-primary system such as PostgreSQL, every write goes through one server, and transactions or joins that span machines are slow and hard to get right. That is why databases usually scale vertically first and horizontally only in stages. A sensible order:
| Step | What it solves | What it costs |
|---|---|---|
| 1. Fix queries and indexes | Slow queries that look like a capacity problem | Engineering time |
| 2. Scale the instance up | CPU and memory headroom | A restart or failover, and a higher hourly price |
| 3. Pool connections | Too many connections, and the cost of one server process per connection | A pooler to run and monitor |
| 4. Add read replicas | Read-heavy traffic such as dashboards and reports | Replication lag, and routing reads correctly |
| 5. Cache hot reads | Repeated queries, such as question papers and leaderboards | Invalidation logic |
| 6. Partition big tables | Very large tables, such as test attempts and logs | Schema and query design |
| 7. Shard across servers | More writes than one primary can take | Cross-shard queries, rebalancing and far more operational work |
Steps 4 and 7 are the database's forms of horizontal scaling. Replicas spread reads: PostgreSQL's hot standby lets a replica answer read-only queries while it replays changes from the primary. Sharding spreads writes, either in your application or through an extension such as Citus, which distributes Postgres tables across several nodes. Our guides to read replicas and sharding vs partitioning go deeper into both.
Cost comparison
You might expect big machines to cost more per unit of power. On AWS on-demand pricing, they don't, within an instance family. In the Mumbai region at the time of writing, m7i.large costs about $0.106 an hour for 2 vCPUs and m7i.48xlarge about $10.18 an hour for 192: exactly 96 times the price for 96 times the vCPUs and memory. The real cost difference comes from how long you run the capacity, and how much spare you keep.
Take an illustrative exam platform whose app tier needs 32 vCPUs during a two-hour evening peak and 8 vCPUs for the rest of the day, using Mumbai on-demand Linux prices and ignoring storage, data transfer and discounts:
| Plan | What runs | Cost per day |
|---|---|---|
| Vertical, one server | One m7i.8xlarge (32 vCPUs) for 24 hours | About $40.7, with no redundancy |
| Vertical, with a standby | Two m7i.8xlarge for 24 hours | About $81.4 |
| Horizontal | Two m7i.xlarge (8 vCPUs) all day in two zones, plus six more for the two-hour peak | About $12.7 |
The horizontal plan is cheaper not because its servers are cheaper, but because it follows the load curve and gets redundancy from running more than one server anyway. The plan only works if the extra capacity arrives before the peak, which is why exam-day platforms schedule it rather than waiting for autoscaling to react. For databases the sum works differently: a primary can't follow a daily curve, so you pay for peak size and buy efficiency through pooling, caching and replicas instead.
A practical path for a growing platform
- Start small: one app server and a managed database, and scale both up as traffic grows.
- Add a second app server early, for redundancy rather than capacity, and use the move to make the app stateless.
- Scale the app tier out with an autoscaling group, a CDN for static files and video, and queue workers for slow jobs. Add a connection pooler at the same time.
- Add read replicas and caching when reads dominate the database's load.
- Partition the biggest tables and move heavy workloads such as analytics or search to their own systems.
- Shard only when you must: when one primary, at the largest sensible size, still can't keep up with writes.
Our guide to handling 100,000 concurrent users shows how these layers fit together on a busy exam day.
Key takeaways
- Vertical scaling is simple but has a ceiling, a single point of failure and downtime to resize.
- Horizontal scaling has no practical ceiling and survives failures, but only if the app keeps no state on its servers.
- On AWS on-demand pricing, cost rises in proportion to size within an instance family; savings come from following the load curve.
- Databases scale up first, then spread reads with replicas, and spread writes with sharding only as a last resort.
- Most growing platforms scale app servers out and the database up.
Frequently asked questions
What is horizontal scaling and vertical scaling?
Vertical scaling adds resources to one machine: more CPU cores, more memory or faster storage. Horizontal scaling adds more machines and uses a load balancer to share the work between them. Vertical is simpler and needs no code changes, but it has a size limit and usually a restart. Horizontal can keep growing and survive a failed server, but the application must store sessions, files and jobs somewhere shared.
What is horizontal scaling in database?
In databases, horizontal scaling means spreading data or queries across several servers. Read replicas copy the whole database and serve read-only queries, which scales reads but not writes. Sharding splits the data itself, for example by institute or by user, so each server holds and writes only part of it. Sharding scales writes but makes cross-shard queries, transactions and rebalancing much harder, so most teams try bigger servers, pooling and replicas first.
What is horizontal scaling in system design?
In system design, horizontal scaling means designing a service so capacity grows by adding identical instances rather than by buying a bigger one. That requires stateless servers behind a load balancer, shared stores for sessions and files, idempotent operations that can be retried safely, and a data layer that won't collapse as instance counts grow. Interviewers usually expect you to explain which parts scale out easily and which, like the database, need replicas, caching or sharding.