Nginx proxy_cache and microcaching: a practical guide
Set up Nginx proxy_cache, microcache dynamic pages for one second, bypass logged-in users, and stop pile-ups with cache locks and stale-while-updating.
On this page 8 sections
Nginx's proxy_cache stores responses from your application servers on the proxy's disk and serves repeat requests without touching the application. Microcaching is the same feature with a lifetime of about one second: long enough to turn a burst of identical requests into a single render, short enough that nobody notices the staleness. Add a cache lock and stale-while-updating, and a public page that would swamp your app servers at 10 a.m. costs about one render a second on each proxy.
Reverse proxy caching in one diagram
Student → CDN → Nginx (proxy_cache) → Gunicorn/Django → Redis, Postgres
│
├─ HIT: answered from Nginx's cache
└─ MISS: one request goes upstream; the response is stored
A reverse proxy receives requests on behalf of your application servers. Because it already sees every request and response, it is a natural place to cache whole responses: the key is built from the request, the value is the response, and your Python code never runs on a hit. It sits one layer behind the CDN in the stack described in types of caching. Because it is right next to the application, each refill is one short local request, which makes caching a page for just a second or two worthwhile.
A minimal proxy_cache setup
Two directives do the work: proxy_cache_path defines the cache, and proxy_cache switches it on for a location.
proxy_cache_path /var/cache/nginx/app levels=1:2 keys_zone=app:16m
max_size=2g inactive=10m use_temp_path=off;
server {
location / {
proxy_pass http://django_app;
proxy_cache app;
proxy_cache_key $scheme$host$request_uri;
proxy_cache_valid 200 1s;
add_header X-Cache-Status $upstream_cache_status;
}
}
- keys_zone=app:16m is shared memory for keys and metadata. The Nginx docs put it at about 8,000 keys per megabyte, so 16 MB covers roughly 1,28,000 cached URLs. The responses themselves live on disk, capped by max_size.
- inactive=10m deletes entries nobody has requested for ten minutes, whether or not they are still fresh. Ten minutes is also the default.
- proxy_cache_key defaults to $scheme$proxy_host$request_uri, which uses the upstream's name rather than the site's host name. If several sites proxy to the same upstream, include $host as above or their pages can mix.
- proxy_cache_valid 200 1s applies only when the application sends no caching headers of its own. Headers from the app (X-Accel-Expires, Expires or Cache-Control) take priority.
- X-Cache-Status exposes $upstream_cache_status (HIT, MISS, EXPIRED, STALE, UPDATING, REVALIDATED or BYPASS) so you can see what is happening.
Nginx is also cautious by default. Its caching guide notes that it won't cache responses marked private, no-cache or no-store, responses that set a cookie, or anything other than GET and HEAD requests, and nothing is cached if proxy_buffering is off. You can override these with proxy_ignore_headers, but think twice before ignoring Set-Cookie: that header is often the only sign a page is personal.
Microcaching dynamic pages
Microcaching works because traffic spikes are made of identical requests. Take an illustrative results day: the public "toppers" page for a mock test gets 3,000 requests a second at 10 a.m., and Django takes 120 ms to render it with two database queries.
- Without a cache, Little's Law says you need 3,000 × 0.12 = 360 requests in progress at once, so 360 busy workers, plus 6,000 queries a second on the database.
- With a one-second microcache and a cache lock, each Nginx server renders the page about once a second. With three Nginx servers, that is about three renders and six queries a second.
- The cost is that a student may see the page up to about a second old, plus the time the refresh takes.
Microcaching suits pages that are public, expensive and busy: the course catalogue, a test's instructions page, results summaries, announcements and public API responses. It doesn't suit anything that shows who the visitor is.
Bypassing the cache for logged-in users
Most pages on a learning platform are personal once a student logs in. The simplest rule is: if the session cookie is present, skip the cache in both directions.
map $cookie_sessionid $skip_cache {
default 1;
"" 0;
}
# inside the location block
proxy_cache_bypass $skip_cache; # don't answer from the cache
proxy_no_cache $skip_cache; # don't store the response
proxy_cache_bypass alone would still store a logged-in student's page for others to receive; proxy_no_cache stops that. The map goes in the http block, and sessionid is Django's default session cookie name.
Django adds its own wrinkle. When a view touches the session, or renders a form with a CSRF token, Django adds Vary: Cookie to the response, and the CSRF middleware also sets a cookie. Nginx honours Vary by storing a separate copy for each distinct Cookie header, and it won't cache a response with Set-Cookie at all. So a "public" page that quietly reads the session or includes a login form won't be microcached. Fix the page: keep public views free of session access and forms, or load the personal parts through a separate request. Our guide to Django caching covers the application side.
Stale content and cache locks
A one-second cache expires sixty times a minute, and each expiry is a chance for a pile-up. Three directives prevent it:
proxy_cache_lock on;
proxy_cache_use_stale updating error timeout http_500 http_502 http_503 http_504;
proxy_cache_background_update on;
- proxy_cache_lock lets only one request fetch a missing entry from the application; the others wait for it to land in the cache. They wait up to proxy_cache_lock_timeout (5 seconds by default), after which a request goes upstream but its response isn't cached.
- proxy_cache_use_stale updating serves the expired copy while one request refreshes it, so an entry that already exists never makes anyone wait. The error, timeout and http_5xx options also keep serving the last good copy if the application is down.
- proxy_cache_background_update performs that refresh as a background subrequest, so even the request that triggers it gets the stale copy immediately.
The lock covers the first fill; stale-while-updating covers every refresh after that. Together they stop the cache stampede that a short TTL would otherwise invite. For pages cached longer than a few seconds, add proxy_cache_revalidate on, so Nginx refreshes with a conditional request and an unchanged page costs a 304 instead of a full render.
Nginx vs Varnish vs load balancer caching
| Option | How it caches | Good for | Watch out for |
|---|---|---|---|
| Nginx proxy_cache | Responses on disk, keys in shared memory | Microcaching, serving stale during outages, when Nginx is already your proxy | Selective purging is part of the commercial NGINX Plus, not open-source Nginx |
| Varnish | Objects, usually in memory, configured with VCL | Complex cache logic, purging, request coalescing out of the box | Another component to run; VCL has a learning curve |
| HAProxy cache | Small objects in RAM | Icons, stylesheets and other small files | Caches only 200 responses with an explicit lifetime or validator |
| Managed layer 7 load balancer | Usually none | Routing, TLS termination, health checks | AWS's Application Load Balancer has no response cache; Google Cloud's caches only with Cloud CDN switched on |
| Layer 4 load balancer | None; it never sees HTTP | Raw TCP or UDP balancing | Can't cache, by design |
Varnish's grace mode is its equivalent of stale-while-updating, and it coalesces simultaneous requests for the same object automatically. If you already run Nginx in front of Django, start with proxy_cache. Reach for Varnish when cache rules become logic rather than configuration. Either way, the load balancer's job stays routing; our comparison of L4 vs L7 load balancers covers that side.
Key takeaways
- proxy_cache_path defines the cache, proxy_cache enables it, and $upstream_cache_status shows whether it's working.
- A one-second microcache turns thousands of identical requests a second into roughly one render per second per proxy.
- Bypass and don't store for logged-in users; use proxy_cache_bypass and proxy_no_cache together.
- Nginx won't cache responses with Set-Cookie or private, no-cache or no-store, so keep public pages free of sessions and forms.
- Turn on proxy_cache_lock, proxy_cache_use_stale updating and proxy_cache_background_update to stop pile-ups at every expiry.
Frequently asked questions
What is reverse proxy in Nginx?
A reverse proxy is a server that receives client requests on behalf of other servers and forwards them. In Nginx, the proxy_pass directive does this: Nginx accepts the student's HTTPS request, passes it to an application server such as Gunicorn, and returns the response. Along the way it can terminate TLS, compress, buffer slow clients, cache responses with proxy_cache and spread requests across several upstream servers.
What is reverse proxy vs load balancer?
A reverse proxy is defined by where it sits: in front of your servers, speaking for them. A load balancer is defined by what it does: spreading requests across several servers and avoiding unhealthy ones. The two overlap heavily. Nginx and HAProxy are reverse proxies that balance load, while a layer 4 load balancer only forwards connections and never reads the HTTP requests inside them.
Can a load balancer act as a reverse proxy?
A layer 7 load balancer already is one: it terminates the client's connection, reads the HTTP request and opens its own connection to a backend. AWS's Application Load Balancer and Nginx with an upstream block both work this way. What a load balancer usually doesn't do is cache; for that you need a proxy with a cache, such as Nginx or Varnish, or a CDN in front.