L4 vs L7 load balancers: differences and when to use each

What layer 4 and layer 7 load balancers can see and do, ALB vs NLB on AWS, load balancing algorithms, and a typical layout for a learning platform.

8 min read
On this page 8 sections
  1. What layer 4 and layer 7 mean
  2. What an L7 balancer can do
  3. When L4 is the better choice
  4. ALB vs NLB
  5. Load balancing algorithms
  6. A typical layout for a learning platform
  7. Key takeaways
  8. Frequently asked questions

A layer 4 (L4) load balancer routes TCP or UDP connections using only addresses and ports, without reading the traffic inside them; a layer 7 (L7) load balancer terminates HTTP, reads each request and can route by host, path, header or cookie. Use L7 for websites and APIs that need TLS termination, path-based routing and per-request balancing, and L4 for raw throughput, non-HTTP protocols, static IP addresses or TLS that must pass through untouched. On AWS, that is the difference between an Application Load Balancer (ALB) and a Network Load Balancer (NLB).

What layer 4 and layer 7 mean

The names come from the OSI networking model. Layer 4 is the transport layer, where TCP and UDP move data between addresses and ports. Layer 7 is the application layer, where HTTP requests, headers and cookies live. The layer a balancer works at decides what it can see, and therefore what it can decide:

What the balancer can useL4L7
Source and destination IP addresses and portsYesYes
TLS certificate and server nameOnly if it terminates TLSYes; it terminates TLS
HTTP method, host, path, headers and cookiesNoYes
What it balancesConnectionsIndividual requests

The last row has a consequence that surprises people. HTTP/2 and gRPC carry many requests over one long-lived connection. An L4 balancer sends that whole connection to one server, and AWS's NLB documentation says each TCP connection goes to a single target for its lifetime. A few busy clients can therefore overload one server while others sit idle. An L7 balancer spreads the individual requests instead.

What an L7 balancer can do

Because it understands HTTP, an L7 balancer is a full reverse proxy, with the powers that come with it:

  • Terminate TLS in one place, so certificates are managed centrally rather than on every server. Our guide to how SSL/TLS works explains what happens during that handshake.

  • Route by host and path, for example /api/ to application servers, /ws/ to WebSocket servers and admin.example.com to an internal service. ALB rules can match on host, path, HTTP headers, request method, query string and source IP.

  • Route by header or cookie, which enables canary releases that send a small share of users to a new version.

  • Handle protocol details such as WebSocket upgrades, HTTP/2 and gRPC, redirects from HTTP to HTTPS, and fixed error pages.

  • Check health at the application level by calling an HTTP path on each server, and pass the client's IP address to servers in the X-Forwarded-For header.

  • Protect and cache: web application firewalls and rate limiting sit naturally at L7, and proxies such as Nginx can cache responses; see Nginx proxy caching.

The price is work. An L7 balancer decrypts and parses every request, holds your TLS keys, and becomes part of your application's behaviour, so its timeouts and limits matter.

When L4 is the better choice

  • The traffic isn't HTTP. Database connections to PostgreSQL or PgBouncer, Redis, MQTT, custom TCP protocols and UDP traffic all need L4.

  • TLS must reach your servers intact, for example when the application itself verifies client certificates.

  • Partners need fixed IP addresses to allow-list. An NLB has one static IP address per Availability Zone, and an internet-facing NLB can use an Elastic IP address in each zone.

  • Throughput is extreme or spiky. AWS says an NLB can handle millions of requests per second.

  • You want the client's IP without HTTP headers. An NLB can preserve the source address, or pass it in a proxy protocol version 2 header.

  • You run your own L7 tier. A common pattern puts an L4 balancer in front of a fleet of Nginx or Envoy proxies that do the HTTP work.

ALB vs NLB

FeatureApplication Load BalancerNetwork Load Balancer
Layer74
ProtocolsHTTP and HTTPS, including HTTP/2, gRPC and WebSocketsTCP, UDP, TLS and QUIC
How a target is chosenListener rules pick a target group; within it, round robin by default, or least outstanding requests, or weighted randomA hash of the connection's protocol, addresses and ports (plus the TCP sequence number)
Idle timeout60 seconds by default, adjustable from 1 to 4,000350 seconds for TCP, adjustable from 60 to 6,000; 350 fixed for TLS listeners; 120 for UDP
Static IP addressesNoYes, one per zone
Client IP at the serverX-Forwarded-For headerPreserved, or sent with proxy protocol v2
Cross-zone balancingAlways onOff by default
WebSocketsNative; each connection stays with the target that accepted the upgradeCarried as ordinary TCP connections

The idle timeouts matter more than they look. AWS recommends setting your application's own keep-alive timeout higher than the ALB's, or the ALB may send a request down a connection the server has just closed and return a 502 error to the user. For WebSockets, send heartbeats more often than the idle timeout; our guide to scaling WebSockets covers the details.

Load balancing algorithms

AlgorithmHow it picks a serverGood forExamples
Round robinEach server in turnSimilar requests on similar serversALB and Nginx default
WeightedIn turn, in proportion to weightsMixed server sizes and gradual rolloutsNginx weight; ALB weighted target groups
Least connectionsThe server with the fewest open connectionsLong-lived connections such as WebSocketsNginx least_conn
Least outstanding requestsThe server with the fewest requests in progressRequests of very different cost, such as a test submission next to a static pageALB
Random with two choicesPick two servers at random and use the less loaded oneLarge fleets behind several balancersNginx random two least_conn
HashThe same key, such as an IP address or URL, always maps to the same serverCache locality and simple affinityNginx ip_hash and hash
Flow hashA hash of each connection's addresses and portsL4 balancing at very high ratesNLB

Round robin assumes every request costs the same. On an exam platform they don't: fetching a cached question is cheap, while submitting and scoring an attempt is not. Least-outstanding-requests or least-connections balancing keeps one server from queueing behind a few expensive requests while its neighbours are idle.

A typical layout for a learning platform

Every platform arranges these pieces differently. As an illustration, a common arrangement looks like this:

  1. A CDN at the edge serves the app bundle, images and video segments, so most bytes never reach your balancers.

  2. An L7 balancer terminates TLS and routes by path: /api/ to the application servers, /ws/ to WebSocket servers with least-connections balancing and a longer idle timeout, and admin paths only from the office network.

  3. Internal L4 connections carry the app servers' non-HTTP traffic, such as connections to the database's connection pooler and to the cache.

  4. Live-class media takes its own path: HLS segments through the CDN, or WebRTC media over UDP to media servers.

  5. Timeouts are lined up: application keep-alive longer than the balancer's idle timeout, and the balancer's idle timeout longer than the WebSocket heartbeat interval.

  6. Deployments drain connections: a server leaving the target group keeps its in-flight requests for the deregistration delay, 300 seconds by default on AWS, while new requests go elsewhere.

For how the balancer fits with caching, queues and databases on a heavy day, see handling 100,000 concurrent users. AWS's NLB overview and the Nginx load balancing guide are good primary references.

Key takeaways

  • L4 balances connections using addresses and ports; L7 reads and balances individual HTTP requests.

  • Choose L7 for web apps and APIs: TLS termination, path routing, WebSocket upgrades and per-request decisions.

  • Choose L4 for non-HTTP traffic, TLS passthrough, static IPs and very high throughput.

  • HTTP/2 and gRPC connections behind an L4 balancer can load servers unevenly.

  • Match the algorithm to the traffic, and line up idle timeouts from client to server.

Frequently asked questions

Is load balancer a reverse proxy?

An L7 load balancer is a reverse proxy: it accepts the client's connection, reads the request, and opens its own connection to a chosen server. L4 balancers vary. Some also proxy, accepting a TCP connection and opening a new one, as Nginx's stream module and HAProxy in TCP mode do. Others forward connections or packets while keeping the client's address, as AWS's Network Load Balancer can.

Is load balancer a single point of failure?

It can be, if you run one balancer on one machine. Managed load balancers avoid this by running nodes in several Availability Zones and publishing several addresses in DNS. If you run your own, use at least two instances with a floating IP address that moves on failure, or several DNS records with health checks. Also treat the balancer's configuration and certificates as critical, because one bad change affects every request.

How do load balancers work?

A load balancer listens on an address and port, and forwards each incoming connection or request to one server in a pool. It picks the server with an algorithm such as round robin or least connections, and skips servers that fail regular health checks. An L4 balancer decides using addresses and ports only; an L7 balancer can also use the host, path, headers and cookies of each HTTP request.

Share this article

Looking for something else?

Talk to Us