Kubernetes HPA: horizontal pod autoscaling explained
How the HorizontalPodAutoscaler calculates replicas, why CPU is often the wrong metric, custom metrics, HPA vs VPA vs KEDA, node autoscaling and tuning.
On this page 8 sections
The Kubernetes HorizontalPodAutoscaler (HPA) is a built-in controller that changes the number of replicas in a Deployment or StatefulSet to keep a metric close to a target. By default it checks every 15 seconds, calculates desiredReplicas = ceil(currentReplicas × currentMetric ÷ target), ignores differences within 10%, scales up at once and scales down only after a five-minute stabilisation window. It adds pods, not machines, so it needs a node autoscaler underneath it and a metric that genuinely tracks load.
How the HPA decides replica counts
The HPA is a control loop inside the kube-controller-manager. Once per sync period (15 seconds unless the cluster sets --horizontal-pod-autoscaler-sync-period), it reads the metric for every pod the target selects, averages it, and applies one formula:
desiredReplicas = ceil(currentReplicas × currentMetricValue ÷ desiredMetricValue)
Three worked examples with a CPU target of 60%:
- Scale up: 6 pods averaging 90% CPU gives ceil(6 × 90 ÷ 60) = 9 pods.
- No change: 10 pods at 64% gives a ratio of 1.07, inside the default 0.1 tolerance, so nothing happens.
- Scale down, later: 10 pods at 30% gives 5 pods, but the HPA uses the highest recommendation from the last five minutes, so the drop only happens once load has stayed low that long.
Some details matter in practice:
- Utilisation is relative to requests. 60% CPU means 60% of the CPU the pod requested. If a container has no CPU request, the pod's utilisation is undefined and the HPA takes no action for that metric.
- Starting pods are handled cautiously. Pods that aren't ready yet, or have no metrics, are set aside at first. The HPA then recalculates conservatively: pods with missing metrics count as using 100% of the target when it would scale down and 0% when it would scale up, and not-yet-ready pods count as idle when it would scale up. Both rules damp the move.
- Several metrics mean the biggest answer wins. The HPA works out a replica count for each metric and uses the largest.
The default scaling behaviour, which you can override per HPA:
| Direction | Stabilisation window | Maximum rate |
|---|---|---|
| Scale up | None; acts immediately | The larger of 100% of current pods or 4 pods, every 15 seconds |
| Scale down | 300 seconds | Up to 100% of current pods every 15 seconds, once the window allows it |
A minimal autoscaling/v2 HPA for a web Deployment, with a gentler scale-down of 10% of pods per minute:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: web}
spec:
scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: web}
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target: {type: Utilization, averageUtilization: 60}
behavior:
scaleDown:
policies: [{type: Percent, value: 10, periodSeconds: 60}]
Choosing the right metric
CPU is the default because it needs nothing beyond Metrics Server, but it is often the wrong signal for a web app. A Django app under Gunicorn's sync workers spends much of each request waiting on the database, Redis or an SMS API. When all workers are busy and requests are queueing, CPU can still read 30%, so the HPA sees no reason to act. And when the real problem is a slow database, adding pods only adds database connections.
| Metric | Works well for | Watch out for |
|---|---|---|
| CPU utilisation | CPU-bound services such as rendering, encoding or heavy JSON work | Unrealistic CPU requests, and I/O-bound apps that saturate without using much CPU |
| Memory utilisation | Workloads whose memory use rises and falls with load | In many web apps memory stays high after load falls, so the HPA struggles to scale back down |
| Requests per second per pod | Web and API tiers with a known, load-tested capacity per pod | Requests that vary widely in cost |
| In-flight requests or busy workers per pod | I/O-bound apps; it measures saturation directly | Needs an exporter and an adapter |
| Queue length per worker | Background workers such as Celery | Needs an external metric; KEDA makes this easiest |
A useful rule: scale on the resource your app runs out of first. For most I/O-bound web apps that is worker slots, not CPU.
Custom and external metrics
The HPA reads metrics from three APIs. metrics.k8s.io carries CPU and memory, usually served by Metrics Server, which collects them every 15 seconds and is meant only for autoscaling. custom.metrics.k8s.io carries metrics about Kubernetes objects, such as requests per pod, and is served by an adapter; the Prometheus Adapter, for example, answers it from Prometheus queries. external.metrics.k8s.io carries metrics from outside the cluster, such as the length of a queue.
With an adapter in place, a Pods metric lets you target an average value per pod. Here the target is 12 requests in flight per pod; both the metric name and the number are examples to replace with your own:
metrics:
- type: Pods
pods:
metric:
name: http_requests_in_flight
target:
type: AverageValue
averageValue: "12"
The Kubernetes HPA documentation covers Object and External metric types, which work the same way for a single value such as queue length.
HPA vs VPA vs KEDA
| Tool | What it changes | How you get it | Best for |
|---|---|---|---|
| HPA | The number of pods | Built into Kubernetes | Stateless web and API services |
| VPA | Each pod's CPU and memory requests and limits | A separate add-on installed as custom resources | Right-sizing requests for steady workloads |
| KEDA | Drives an HPA from event sources, and scales between zero and one pod itself | A separate add-on; a CNCF graduated project | Queue workers, schedule-based scaling and scaling to zero |
They can work together, with care. The VPA project says not to use VPA and HPA on the same CPU or memory metric, because each would keep reacting to the other's changes: one resizes pods while the other counts them. A safe pattern is VPA in its recommendation-only Off mode to find realistic requests, and HPA on requests or saturation. VPA's other modes apply changes when pods are created, by evicting pods, or by resizing them in place where the cluster supports it; in-place pod resizing is stable from Kubernetes 1.35.
KEDA suits queue workers. It creates and manages an HPA for you, feeds it metrics from sources such as Redis lists, RabbitMQ, Kafka, SQS or Prometheus, and handles the step from zero pods to one itself. Its cron trigger can also hold a minimum number of pods between two times of day. Scale to zero isn't possible with CPU or memory triggers, because there is nothing to measure when no pods are running. Plain HPA can now scale to zero too, but only on object or external metrics: the HPAScaleToZero feature is on by default from Kubernetes 1.37.
Node scaling underneath
The HPA only asks for pods. If the nodes have no room, new pods sit in Pending until a node autoscaler adds capacity. Kubernetes SIG Autoscaling sponsors two:
- Cluster Autoscaler adds and removes nodes in node groups you configure in advance. By default it looks for unschedulable pods every 10 seconds and removes a node once it has been unneeded for 10 minutes.
- Karpenter picks and launches individual machines to fit pending pods, within the limits of NodePools you define, and also manages node lifetimes and upgrades. It has providers for AWS and Azure.
Both work from pods' resource requests, not their real usage, which is another reason to set requests carefully. Timing matters too. The Cluster Autoscaler FAQ puts the whole chain, from rising load to pods running on new nodes, at usually about five minutes, most of it spent provisioning the node.
To hide that delay, keep a little spare room. The Cluster Autoscaler FAQ describes overprovisioning with low-priority placeholder pods: a Deployment of pause pods with a PriorityClass value of -10 reserves space, real pods preempt them the moment they need it, and the evicted placeholders go Pending and trigger a new node in the background.
Tuning behaviour and stabilisation
- Leave scale-up fast. The default already adds pods quickly; the real delay is pod start-up and node provisioning.
- Slow down scale-down when traffic comes in waves, such as a test-start rush followed by a submission rush, so capacity isn't removed between them.
- Make readiness honest. A startup or readiness probe that passes only when the app can serve stops warming pods from skewing the metric.
- Set tolerance per HPA where needed. From Kubernetes 1.37 you can set a tolerance per HPA and per direction; older clusters use the cluster-wide 10% default (1.33 to 1.36 offered it behind a feature gate).
- Remove
replicasfrom the Deployment manifest once an HPA manages it. Otherwise everykubectl applyresets the replica count. - Raise
minReplicasbefore known peaks instead of hoping the HPA reacts in time; see scheduled and predictive scaling. - Budget database connections.
maxReplicasmultiplied by each pod's pool size must fit what the database can take; see sizing a connection pool.
The same ideas apply outside Kubernetes; how autoscaling works covers the EC2 Auto Scaling version.
Key takeaways
- The HPA scales pod counts in proportion to how far a metric is from its target, every 15 seconds by default.
- Scale-up is immediate; scale-down waits for a five-minute window unless you change it.
- CPU is only a good metric for CPU-bound services; I/O-bound web apps need a request or saturation metric.
- Don't let VPA and HPA act on the same resource metric; use KEDA for queue workers and scale to zero.
- Pods need nodes: pair the HPA with Cluster Autoscaler or Karpenter, and keep spare room for spikes.
Frequently asked questions
What is HPA in Kubernetes?
HPA stands for HorizontalPodAutoscaler, a built-in Kubernetes API resource and controller. It watches a metric for the pods of a Deployment or StatefulSet, such as CPU utilisation or requests per pod, and changes the replica count to keep that metric near a target. It works between the minReplicas and maxReplicas you set, scales up quickly, scales down cautiously, and can't scale objects such as DaemonSets.
What is HPA and VPA in Kubernetes?
HPA scales horizontally by changing how many pods run. VPA, the VerticalPodAutoscaler, scales vertically by changing each pod's CPU and memory requests based on observed usage. HPA is built in, while VPA is installed separately. Don't let both act on the same CPU or memory metric; a common pattern is VPA in recommendation-only mode to right-size requests, with HPA scaling on requests or a custom metric.
What is HPA in K8s?
K8s is shorthand for Kubernetes, so HPA in K8s is the same HorizontalPodAutoscaler. You define it in YAML with the autoscaling/v2 API, or create a basic one with kubectl autoscale. kubectl get hpa shows each autoscaler's current and target metric and replica counts, and kubectl describe hpa lists recent scaling events and conditions, which is the first place to look when it isn't scaling.