Kubernetes HPA: horizontal pod autoscaling explained

How the HorizontalPodAutoscaler calculates replicas, why CPU is often the wrong metric, custom metrics, HPA vs VPA vs KEDA, node autoscaling and tuning.

9 min read
On this page 8 sections
  1. How the HPA decides replica counts
  2. Choosing the right metric
  3. Custom and external metrics
  4. HPA vs VPA vs KEDA
  5. Node scaling underneath
  6. Tuning behaviour and stabilisation
  7. Key takeaways
  8. Frequently asked questions

The Kubernetes HorizontalPodAutoscaler (HPA) is a built-in controller that changes the number of replicas in a Deployment or StatefulSet to keep a metric close to a target. By default it checks every 15 seconds, calculates desiredReplicas = ceil(currentReplicas × currentMetric ÷ target), ignores differences within 10%, scales up at once and scales down only after a five-minute stabilisation window. It adds pods, not machines, so it needs a node autoscaler underneath it and a metric that genuinely tracks load.

How the HPA decides replica counts

The HPA is a control loop inside the kube-controller-manager. Once per sync period (15 seconds unless the cluster sets --horizontal-pod-autoscaler-sync-period), it reads the metric for every pod the target selects, averages it, and applies one formula:

desiredReplicas = ceil(currentReplicas × currentMetricValue ÷ desiredMetricValue)

Three worked examples with a CPU target of 60%:

  • Scale up: 6 pods averaging 90% CPU gives ceil(6 × 90 ÷ 60) = 9 pods.

  • No change: 10 pods at 64% gives a ratio of 1.07, inside the default 0.1 tolerance, so nothing happens.

  • Scale down, later: 10 pods at 30% gives 5 pods, but the HPA uses the highest recommendation from the last five minutes, so the drop only happens once load has stayed low that long.

Some details matter in practice:

  • Utilisation is relative to requests. 60% CPU means 60% of the CPU the pod requested. If a container has no CPU request, the pod's utilisation is undefined and the HPA takes no action for that metric.

  • Starting pods are handled cautiously. Pods that aren't ready yet, or have no metrics, are set aside at first. The HPA then recalculates conservatively: pods with missing metrics count as using 100% of the target when it would scale down and 0% when it would scale up, and not-yet-ready pods count as idle when it would scale up. Both rules damp the move.

  • Several metrics mean the biggest answer wins. The HPA works out a replica count for each metric and uses the largest.

The default scaling behaviour, which you can override per HPA:

DirectionStabilisation windowMaximum rate
Scale upNone; acts immediatelyThe larger of 100% of current pods or 4 pods, every 15 seconds
Scale down300 secondsUp to 100% of current pods every 15 seconds, once the window allows it

A minimal autoscaling/v2 HPA for a web Deployment, with a gentler scale-down of 10% of pods per minute:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: web}
spec:
  scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: web}
  minReplicas: 3
  maxReplicas: 30
  metrics:
  - type: Resource
    resource:
      name: cpu
      target: {type: Utilization, averageUtilization: 60}
  behavior:
    scaleDown:
      policies: [{type: Percent, value: 10, periodSeconds: 60}]

Choosing the right metric

CPU is the default because it needs nothing beyond Metrics Server, but it is often the wrong signal for a web app. A Django app under Gunicorn's sync workers spends much of each request waiting on the database, Redis or an SMS API. When all workers are busy and requests are queueing, CPU can still read 30%, so the HPA sees no reason to act. And when the real problem is a slow database, adding pods only adds database connections.

MetricWorks well forWatch out for
CPU utilisationCPU-bound services such as rendering, encoding or heavy JSON workUnrealistic CPU requests, and I/O-bound apps that saturate without using much CPU
Memory utilisationWorkloads whose memory use rises and falls with loadIn many web apps memory stays high after load falls, so the HPA struggles to scale back down
Requests per second per podWeb and API tiers with a known, load-tested capacity per podRequests that vary widely in cost
In-flight requests or busy workers per podI/O-bound apps; it measures saturation directlyNeeds an exporter and an adapter
Queue length per workerBackground workers such as CeleryNeeds an external metric; KEDA makes this easiest

A useful rule: scale on the resource your app runs out of first. For most I/O-bound web apps that is worker slots, not CPU.

Custom and external metrics

The HPA reads metrics from three APIs. metrics.k8s.io carries CPU and memory, usually served by Metrics Server, which collects them every 15 seconds and is meant only for autoscaling. custom.metrics.k8s.io carries metrics about Kubernetes objects, such as requests per pod, and is served by an adapter; the Prometheus Adapter, for example, answers it from Prometheus queries. external.metrics.k8s.io carries metrics from outside the cluster, such as the length of a queue.

With an adapter in place, a Pods metric lets you target an average value per pod. Here the target is 12 requests in flight per pod; both the metric name and the number are examples to replace with your own:

metrics:
- type: Pods
  pods:
    metric:
      name: http_requests_in_flight
    target:
      type: AverageValue
      averageValue: "12"

The Kubernetes HPA documentation covers Object and External metric types, which work the same way for a single value such as queue length.

HPA vs VPA vs KEDA

ToolWhat it changesHow you get itBest for
HPAThe number of podsBuilt into KubernetesStateless web and API services
VPAEach pod's CPU and memory requests and limitsA separate add-on installed as custom resourcesRight-sizing requests for steady workloads
KEDADrives an HPA from event sources, and scales between zero and one pod itselfA separate add-on; a CNCF graduated projectQueue workers, schedule-based scaling and scaling to zero

They can work together, with care. The VPA project says not to use VPA and HPA on the same CPU or memory metric, because each would keep reacting to the other's changes: one resizes pods while the other counts them. A safe pattern is VPA in its recommendation-only Off mode to find realistic requests, and HPA on requests or saturation. VPA's other modes apply changes when pods are created, by evicting pods, or by resizing them in place where the cluster supports it; in-place pod resizing is stable from Kubernetes 1.35.

KEDA suits queue workers. It creates and manages an HPA for you, feeds it metrics from sources such as Redis lists, RabbitMQ, Kafka, SQS or Prometheus, and handles the step from zero pods to one itself. Its cron trigger can also hold a minimum number of pods between two times of day. Scale to zero isn't possible with CPU or memory triggers, because there is nothing to measure when no pods are running. Plain HPA can now scale to zero too, but only on object or external metrics: the HPAScaleToZero feature is on by default from Kubernetes 1.37.

Node scaling underneath

The HPA only asks for pods. If the nodes have no room, new pods sit in Pending until a node autoscaler adds capacity. Kubernetes SIG Autoscaling sponsors two:

  • Cluster Autoscaler adds and removes nodes in node groups you configure in advance. By default it looks for unschedulable pods every 10 seconds and removes a node once it has been unneeded for 10 minutes.

  • Karpenter picks and launches individual machines to fit pending pods, within the limits of NodePools you define, and also manages node lifetimes and upgrades. It has providers for AWS and Azure.

Both work from pods' resource requests, not their real usage, which is another reason to set requests carefully. Timing matters too. The Cluster Autoscaler FAQ puts the whole chain, from rising load to pods running on new nodes, at usually about five minutes, most of it spent provisioning the node.

To hide that delay, keep a little spare room. The Cluster Autoscaler FAQ describes overprovisioning with low-priority placeholder pods: a Deployment of pause pods with a PriorityClass value of -10 reserves space, real pods preempt them the moment they need it, and the evicted placeholders go Pending and trigger a new node in the background.

Tuning behaviour and stabilisation

  • Leave scale-up fast. The default already adds pods quickly; the real delay is pod start-up and node provisioning.

  • Slow down scale-down when traffic comes in waves, such as a test-start rush followed by a submission rush, so capacity isn't removed between them.

  • Make readiness honest. A startup or readiness probe that passes only when the app can serve stops warming pods from skewing the metric.

  • Set tolerance per HPA where needed. From Kubernetes 1.37 you can set a tolerance per HPA and per direction; older clusters use the cluster-wide 10% default (1.33 to 1.36 offered it behind a feature gate).

  • Remove replicas from the Deployment manifest once an HPA manages it. Otherwise every kubectl apply resets the replica count.

  • Raise minReplicas before known peaks instead of hoping the HPA reacts in time; see scheduled and predictive scaling.

  • Budget database connections. maxReplicas multiplied by each pod's pool size must fit what the database can take; see sizing a connection pool.

The same ideas apply outside Kubernetes; how autoscaling works covers the EC2 Auto Scaling version.

Key takeaways

  • The HPA scales pod counts in proportion to how far a metric is from its target, every 15 seconds by default.

  • Scale-up is immediate; scale-down waits for a five-minute window unless you change it.

  • CPU is only a good metric for CPU-bound services; I/O-bound web apps need a request or saturation metric.

  • Don't let VPA and HPA act on the same resource metric; use KEDA for queue workers and scale to zero.

  • Pods need nodes: pair the HPA with Cluster Autoscaler or Karpenter, and keep spare room for spikes.

Frequently asked questions

What is HPA in Kubernetes?

HPA stands for HorizontalPodAutoscaler, a built-in Kubernetes API resource and controller. It watches a metric for the pods of a Deployment or StatefulSet, such as CPU utilisation or requests per pod, and changes the replica count to keep that metric near a target. It works between the minReplicas and maxReplicas you set, scales up quickly, scales down cautiously, and can't scale objects such as DaemonSets.

What is HPA and VPA in Kubernetes?

HPA scales horizontally by changing how many pods run. VPA, the VerticalPodAutoscaler, scales vertically by changing each pod's CPU and memory requests based on observed usage. HPA is built in, while VPA is installed separately. Don't let both act on the same CPU or memory metric; a common pattern is VPA in recommendation-only mode to right-size requests, with HPA scaling on requests or a custom metric.

What is HPA in K8s?

K8s is shorthand for Kubernetes, so HPA in K8s is the same HorizontalPodAutoscaler. You define it in YAML with the autoscaling/v2 API, or create a basic one with kubectl autoscale. kubectl get hpa shows each autoscaler's current and target metric and replica counts, and kubectl describe hpa lists recent scaling events and conditions, which is the first place to look when it isn't scaling.

Share this article

Looking for something else?

Talk to Us