Types of caching: every layer from browser to database
Browser, CDN, reverse proxy, app memory, Redis and database: what each caching layer stores, who shares it, and how one results-day request moves through them.
On this page 13 sections
- Why cache at all
- Follow one request through the layers
- Browser cache
- CDN edge cache
- Reverse proxy and load balancer caching
- Application and Redis caches
- In-process memory
- Shared cache: Redis or Memcached
- Database-level caching
- What to cache at each layer
- TTLs and invalidation across layers
- Key takeaways
- Frequently asked questions
The main types of caching, in the order a request meets them, are the browser cache, the CDN edge cache, the reverse-proxy cache, the application server's in-process cache, a shared cache such as Redis, and the database's own caches. Each layer keeps a copy of something expensive, whether a file, a page or a query result, closer to where it is needed, so the layers behind it do less work. What differs is who shares the copy, how long it may live and how you get rid of it when the data changes.
Why cache at all
A cache trades a little freshness for a lot of speed. Serving a stored copy is faster than rebuilding it, it takes load off everything behind it, and it saves money, because the costliest resources in a learning platform (database CPU, origin bandwidth and application servers) are exactly the ones a cache shields.
The price is staleness and complexity. Every copy can be out of date, every layer needs a rule for how long copies live, and a mistake can be worse than old data: a shared cache that stores a personal page can show one student's marks to another.
Follow one request through the layers
Take an illustrative case. A mock test's results go live at 10 a.m., and 20,000 students open "My result" in the first five minutes. Each page load needs an HTML page, 25 static files (JavaScript, CSS, fonts and logos), one JSON call for the student's own score and rank, and one for the top-100 leaderboard, which is the same for everyone.
Without caching, that is 20,000 × 28 = 5,60,000 requests to your servers, including 20,000 runs of a heavy leaderboard query. With each layer doing its job, the picture changes:
| Request | Count | Where it is answered | What reaches the database |
|---|---|---|---|
| 25 static files per student | 5,00,000 | Browser cache for returning students, CDN edge for the rest | Nothing |
| HTML page (one shell for everyone; personal data loads separately) | 20,000 | CDN edge, or a one-second reverse-proxy cache | Nothing, apart from a few hundred renders at most |
| Leaderboard JSON (same for everyone) | 20,000 | App server, reading a copy kept in Redis for 30 seconds | About ten heavy queries in five minutes |
| My score and rank (personal) | 20,000 | App server, reading a per-student Redis key or a precomputed table | Up to 20,000 fast indexed lookups |
The database goes from 20,000 heavy aggregations plus 20,000 lookups to about ten aggregations plus cheap lookups, and the application servers handle roughly 40,000 requests instead of 5,60,000. No single layer did that. Each one stopped a different kind of request.
Browser cache
The browser keeps responses on the student's own device, following the Cache-Control header your server sends: max-age for how long a copy stays fresh, no-cache to force a check with the server, no-store to forbid keeping it at all. It is a private cache, which makes it the one layer where personal responses may be stored.
The best candidates are JavaScript, CSS, fonts and images with a content hash in the file name. They can be cached for a year, because a new version gets a new URL. HTML should usually be revalidated on every load or cached only briefly. The catch is that you can't purge a browser: once a phone has been told to keep a file for a day, it will. Service workers add a second browser-side store, the Cache API, which is useful for offline notes but has its own update rules. Our guide to Cache-Control headers gives recipes for each type of asset.
CDN edge cache
A CDN is a shared HTTP cache spread across many cities. Students in Patna and Pune may reach different edge servers, and each edge keeps its own copies of your files, keyed by URL and a few request details. It is the natural home for static assets, video segments and public pages such as the course catalogue.
Shared means careful. Anything a CDN stores can be served to every user whose request matches, so personal pages must be marked private or bypass the CDN entirely. Unlike a browser, though, a CDN can be purged. CDN caching covers cache keys, hit ratios and Cloudflare's cache rules.
Reverse proxy and load balancer caching
In front of your application servers sits a load balancer, a reverse proxy or both. This is where teams most often assume caching that isn't there.
- Layer 4 load balancers forward TCP or UDP connections without reading the HTTP inside them, so they cannot cache responses at all.
- Managed layer 7 load balancers read HTTP, but that doesn't make them caches. AWS's Application Load Balancer routes by host and path, terminates TLS and runs health checks, and on AWS the cache in front of it is CloudFront's job. Google Cloud's external Application Load Balancer caches only when you switch on Cloud CDN for a backend.
- Reverse proxies such as Nginx and Varnish do cache, and they can balance load as well. Nginx stores cached responses on disk. HAProxy also has a cache, but its documentation describes it as a minimal in-memory cache for small objects such as icons and stylesheets.
Because a reverse proxy sits right next to your application, it suits very short-lived caching of dynamic pages. Even one second of Nginx proxy_cache can turn thousands of identical requests a second into one render a second on each proxy.
Application and Redis caches
In-process memory
Inside the application server you can keep values in the process's own memory: a Python dictionary, functools.lru_cache, Django's local-memory backend or cached_property on an object. It is the fastest cache there is, with no network hop, but every process has its own copy. With 8 Gunicorn workers on each of 6 servers, a value lives in 48 separate places, and deleting it in one does nothing to the other 47. Use it for things that rarely change and can be slightly stale, such as configuration or the list of exam categories.
Shared cache: Redis or Memcached
A shared in-memory store gives every app server the same view. This is where most application caching lives: computed results such as a leaderboard or a course outline with lesson counts, rendered fragments, sessions and rate-limit counters. You decide what to store and for how long, usually with the cache-aside pattern: check the cache, and on a miss, compute the value and store it. See Redis caching for eviction and memory sizing, and Django caching for the framework side.
Database-level caching
Databases cache too, mostly without being asked. PostgreSQL keeps recently used 8 kB data pages in shared_buffers, which is 128 MB by default; its documentation suggests about 25% of RAM as a starting point on a dedicated server, and it relies on the operating system's file cache as well. A query whose pages are already in memory skips disk reads, but it still pays for parsing, planning, joining and sorting.
What PostgreSQL doesn't have is a result cache that remembers the answer to a query. For expensive, repeated queries such as leaderboards and reports, you store the result yourself, either in Redis or inside the database as a materialized view that you refresh on a schedule.
What to cache at each layer
In system design terms, these are your caching levels. The table sums up what each one is good for and how you clear it.
| Layer | Shared by | Good for | Keep out | How you clear it |
|---|---|---|---|---|
| Browser | One user | Fingerprinted JS, CSS and fonts; images | Anything you might need to recall in a hurry | You can't; change the URL |
| CDN edge | Everyone near that edge | Static files, video segments, public pages, public API responses | Logged-in pages, personal JSON, payment and OTP responses | Purge by URL, tag or prefix |
| Reverse proxy | Everyone behind that proxy | Public dynamic pages, for one to a few seconds | Anything that varies by user | Short TTL, or purge where supported |
| In-process memory | One worker process | Configuration, small reference lists | Anything that must match across servers | Short TTL or a restart |
| Redis or Memcached | All app servers | Computed results, fragments, sessions, counters | The only copy of important data | Delete or version the key |
| Database | All queries | Hot data pages (automatic), precomputed results in materialized views | Nothing; it is the source of truth | Refresh the view |
TTLs and invalidation across layers
Every cached copy needs an answer to "how long can this be wrong?" Layers make that harder, because staleness stacks. Suppose a leaderboard is cached in Redis for 60 seconds and the page showing it is cached at the CDN for five minutes. A student can see a ranking about six minutes old, because the CDN may have stored a page that was built from Redis data already a minute old. HTTP caches don't add up this way among themselves, since each passes an Age header and the next cache counts it against its own limit. The application cache, however, is invisible to HTTP.
Three habits keep this manageable:
- Set TTLs from the inside out. Decide how stale each kind of data may be, then keep the total across layers within that limit.
- Invalidate in the same order. After a change, clear the application cache first, then the reverse proxy, then the CDN. Purge the CDN first and it may refetch the old value from a layer you haven't cleared yet.
- Version what you can't purge. Browsers can't be purged, so static files get a new URL for each version and HTML stays short-lived.
Two failure modes deserve their own reading. Cache invalidation covers choosing TTLs and purging safely, and cache stampedes covers what happens when a popular entry expires and thousands of requests miss at the same moment.
Key takeaways
- A request can be answered by six layers: the browser, the CDN, a reverse proxy, in-process memory, a shared cache such as Redis, and the database's own memory.
- Only private caches such as the browser may hold personal responses. Shared caches must never serve one user's data to another.
- Load balancers mostly don't cache: layer 4 balancers can't, and many managed layer 7 balancers don't. Reverse proxies such as Nginx and Varnish do.
- PostgreSQL caches data pages, not query results, so store expensive results in Redis or a materialized view.
- Staleness stacks across layers, so set TTLs and invalidate from the inside out.
Frequently asked questions
Why cache is faster than database?
A cache lookup does far less work. Redis or Memcached finds a value by its key in memory and sends it back, while a database has to parse the query, plan it, read pages, apply filters, join tables and sort the result. Databases keep hot pages in memory too, so the biggest saving usually isn't RAM versus disk. It is skipping the computation and returning an answer that was already built.
Is CDN a cache?
Yes. A CDN's core job is to act as a shared HTTP cache spread across many locations, so users get copies from a nearby edge server instead of your origin. It follows the same Cache-Control rules as other shared caches, plus its own settings. Most CDNs also do more than cache: they terminate TLS, absorb attacks, route traffic to healthy origins and can run code at the edge.
How do caches improve performance?
They improve latency, capacity and cost together. Latency drops because the copy comes from somewhere closer or cheaper to reach. Capacity rises because every hit is a request your servers and database never see, so the same hardware survives bigger spikes. Cost falls because origin bandwidth and database CPU are expensive. The size of the gain depends on the hit ratio: at 99% hits, the origin sees one request in a hundred.