L4 vs L7 load balancers: differences and when to use each
What layer 4 and layer 7 load balancers can see and do, ALB vs NLB on AWS, load balancing algorithms, and a typical layout for a learning platform.
On this page 8 sections
A layer 4 (L4) load balancer routes TCP or UDP connections using only addresses and ports, without reading the traffic inside them; a layer 7 (L7) load balancer terminates HTTP, reads each request and can route by host, path, header or cookie. Use L7 for websites and APIs that need TLS termination, path-based routing and per-request balancing, and L4 for raw throughput, non-HTTP protocols, static IP addresses or TLS that must pass through untouched. On AWS, that is the difference between an Application Load Balancer (ALB) and a Network Load Balancer (NLB).
What layer 4 and layer 7 mean
The names come from the OSI networking model. Layer 4 is the transport layer, where TCP and UDP move data between addresses and ports. Layer 7 is the application layer, where HTTP requests, headers and cookies live. The layer a balancer works at decides what it can see, and therefore what it can decide:
| What the balancer can use | L4 | L7 |
|---|---|---|
| Source and destination IP addresses and ports | Yes | Yes |
| TLS certificate and server name | Only if it terminates TLS | Yes; it terminates TLS |
| HTTP method, host, path, headers and cookies | No | Yes |
| What it balances | Connections | Individual requests |
The last row has a consequence that surprises people. HTTP/2 and gRPC carry many requests over one long-lived connection. An L4 balancer sends that whole connection to one server, and AWS's NLB documentation says each TCP connection goes to a single target for its lifetime. A few busy clients can therefore overload one server while others sit idle. An L7 balancer spreads the individual requests instead.
What an L7 balancer can do
Because it understands HTTP, an L7 balancer is a full reverse proxy, with the powers that come with it:
- Terminate TLS in one place, so certificates are managed centrally rather than on every server. Our guide to how SSL/TLS works explains what happens during that handshake.
- Route by host and path, for example
/api/to application servers,/ws/to WebSocket servers andadmin.example.comto an internal service. ALB rules can match on host, path, HTTP headers, request method, query string and source IP. - Route by header or cookie, which enables canary releases that send a small share of users to a new version.
- Handle protocol details such as WebSocket upgrades, HTTP/2 and gRPC, redirects from HTTP to HTTPS, and fixed error pages.
- Check health at the application level by calling an HTTP path on each server, and pass the client's IP address to servers in the
X-Forwarded-Forheader. - Protect and cache: web application firewalls and rate limiting sit naturally at L7, and proxies such as Nginx can cache responses; see Nginx proxy caching.
The price is work. An L7 balancer decrypts and parses every request, holds your TLS keys, and becomes part of your application's behaviour, so its timeouts and limits matter.
When L4 is the better choice
- The traffic isn't HTTP. Database connections to PostgreSQL or PgBouncer, Redis, MQTT, custom TCP protocols and UDP traffic all need L4.
- TLS must reach your servers intact, for example when the application itself verifies client certificates.
- Partners need fixed IP addresses to allow-list. An NLB has one static IP address per Availability Zone, and an internet-facing NLB can use an Elastic IP address in each zone.
- Throughput is extreme or spiky. AWS says an NLB can handle millions of requests per second.
- You want the client's IP without HTTP headers. An NLB can preserve the source address, or pass it in a proxy protocol version 2 header.
- You run your own L7 tier. A common pattern puts an L4 balancer in front of a fleet of Nginx or Envoy proxies that do the HTTP work.
ALB vs NLB
| Feature | Application Load Balancer | Network Load Balancer |
|---|---|---|
| Layer | 7 | 4 |
| Protocols | HTTP and HTTPS, including HTTP/2, gRPC and WebSockets | TCP, UDP, TLS and QUIC |
| How a target is chosen | Listener rules pick a target group; within it, round robin by default, or least outstanding requests, or weighted random | A hash of the connection's protocol, addresses and ports (plus the TCP sequence number) |
| Idle timeout | 60 seconds by default, adjustable from 1 to 4,000 | 350 seconds for TCP, adjustable from 60 to 6,000; 350 fixed for TLS listeners; 120 for UDP |
| Static IP addresses | No | Yes, one per zone |
| Client IP at the server | X-Forwarded-For header | Preserved, or sent with proxy protocol v2 |
| Cross-zone balancing | Always on | Off by default |
| WebSockets | Native; each connection stays with the target that accepted the upgrade | Carried as ordinary TCP connections |
The idle timeouts matter more than they look. AWS recommends setting your application's own keep-alive timeout higher than the ALB's, or the ALB may send a request down a connection the server has just closed and return a 502 error to the user. For WebSockets, send heartbeats more often than the idle timeout; our guide to scaling WebSockets covers the details.
Load balancing algorithms
| Algorithm | How it picks a server | Good for | Examples |
|---|---|---|---|
| Round robin | Each server in turn | Similar requests on similar servers | ALB and Nginx default |
| Weighted | In turn, in proportion to weights | Mixed server sizes and gradual rollouts | Nginx weight; ALB weighted target groups |
| Least connections | The server with the fewest open connections | Long-lived connections such as WebSockets | Nginx least_conn |
| Least outstanding requests | The server with the fewest requests in progress | Requests of very different cost, such as a test submission next to a static page | ALB |
| Random with two choices | Pick two servers at random and use the less loaded one | Large fleets behind several balancers | Nginx random two least_conn |
| Hash | The same key, such as an IP address or URL, always maps to the same server | Cache locality and simple affinity | Nginx ip_hash and hash |
| Flow hash | A hash of each connection's addresses and ports | L4 balancing at very high rates | NLB |
Round robin assumes every request costs the same. On an exam platform they don't: fetching a cached question is cheap, while submitting and scoring an attempt is not. Least-outstanding-requests or least-connections balancing keeps one server from queueing behind a few expensive requests while its neighbours are idle.
A typical layout for a learning platform
Every platform arranges these pieces differently. As an illustration, a common arrangement looks like this:
- A CDN at the edge serves the app bundle, images and video segments, so most bytes never reach your balancers.
- An L7 balancer terminates TLS and routes by path:
/api/to the application servers,/ws/to WebSocket servers with least-connections balancing and a longer idle timeout, and admin paths only from the office network. - Internal L4 connections carry the app servers' non-HTTP traffic, such as connections to the database's connection pooler and to the cache.
- Live-class media takes its own path: HLS segments through the CDN, or WebRTC media over UDP to media servers.
- Timeouts are lined up: application keep-alive longer than the balancer's idle timeout, and the balancer's idle timeout longer than the WebSocket heartbeat interval.
- Deployments drain connections: a server leaving the target group keeps its in-flight requests for the deregistration delay, 300 seconds by default on AWS, while new requests go elsewhere.
For how the balancer fits with caching, queues and databases on a heavy day, see handling 100,000 concurrent users. AWS's NLB overview and the Nginx load balancing guide are good primary references.
Key takeaways
- L4 balances connections using addresses and ports; L7 reads and balances individual HTTP requests.
- Choose L7 for web apps and APIs: TLS termination, path routing, WebSocket upgrades and per-request decisions.
- Choose L4 for non-HTTP traffic, TLS passthrough, static IPs and very high throughput.
- HTTP/2 and gRPC connections behind an L4 balancer can load servers unevenly.
- Match the algorithm to the traffic, and line up idle timeouts from client to server.
Frequently asked questions
Is load balancer a reverse proxy?
An L7 load balancer is a reverse proxy: it accepts the client's connection, reads the request, and opens its own connection to a chosen server. L4 balancers vary. Some also proxy, accepting a TCP connection and opening a new one, as Nginx's stream module and HAProxy in TCP mode do. Others forward connections or packets while keeping the client's address, as AWS's Network Load Balancer can.
Is load balancer a single point of failure?
It can be, if you run one balancer on one machine. Managed load balancers avoid this by running nodes in several Availability Zones and publishing several addresses in DNS. If you run your own, use at least two instances with a floating IP address that moves on failure, or several DNS records with health checks. Also treat the balancer's configuration and certificates as critical, because one bad change affects every request.
How do load balancers work?
A load balancer listens on an address and port, and forwards each incoming connection or request to one server in a pool. It picks the server with an algorithm such as round robin or least connections, and skips servers that fail regular health checks. An L4 balancer decides using addresses and ports only; an L7 balancer can also use the host, path, headers and cookies of each HTTP request.