Types of caching: every layer from browser to database

Browser, CDN, reverse proxy, app memory, Redis and database: what each caching layer stores, who shares it, and how one results-day request moves through them.

10 min read
On this page 13 sections
  1. Why cache at all
  2. Follow one request through the layers
  3. Browser cache
  4. CDN edge cache
  5. Reverse proxy and load balancer caching
  6. Application and Redis caches
  7. In-process memory
  8. Shared cache: Redis or Memcached
  9. Database-level caching
  10. What to cache at each layer
  11. TTLs and invalidation across layers
  12. Key takeaways
  13. Frequently asked questions

The main types of caching, in the order a request meets them, are the browser cache, the CDN edge cache, the reverse-proxy cache, the application server's in-process cache, a shared cache such as Redis, and the database's own caches. Each layer keeps a copy of something expensive, whether a file, a page or a query result, closer to where it is needed, so the layers behind it do less work. What differs is who shares the copy, how long it may live and how you get rid of it when the data changes.

Why cache at all

A cache trades a little freshness for a lot of speed. Serving a stored copy is faster than rebuilding it, it takes load off everything behind it, and it saves money, because the costliest resources in a learning platform (database CPU, origin bandwidth and application servers) are exactly the ones a cache shields.

The price is staleness and complexity. Every copy can be out of date, every layer needs a rule for how long copies live, and a mistake can be worse than old data: a shared cache that stores a personal page can show one student's marks to another.

Follow one request through the layers

Take an illustrative case. A mock test's results go live at 10 a.m., and 20,000 students open "My result" in the first five minutes. Each page load needs an HTML page, 25 static files (JavaScript, CSS, fonts and logos), one JSON call for the student's own score and rank, and one for the top-100 leaderboard, which is the same for everyone.

Without caching, that is 20,000 × 28 = 5,60,000 requests to your servers, including 20,000 runs of a heavy leaderboard query. With each layer doing its job, the picture changes:

RequestCountWhere it is answeredWhat reaches the database
25 static files per student5,00,000Browser cache for returning students, CDN edge for the restNothing
HTML page (one shell for everyone; personal data loads separately)20,000CDN edge, or a one-second reverse-proxy cacheNothing, apart from a few hundred renders at most
Leaderboard JSON (same for everyone)20,000App server, reading a copy kept in Redis for 30 secondsAbout ten heavy queries in five minutes
My score and rank (personal)20,000App server, reading a per-student Redis key or a precomputed tableUp to 20,000 fast indexed lookups

The database goes from 20,000 heavy aggregations plus 20,000 lookups to about ten aggregations plus cheap lookups, and the application servers handle roughly 40,000 requests instead of 5,60,000. No single layer did that. Each one stopped a different kind of request.

Browser cache

The browser keeps responses on the student's own device, following the Cache-Control header your server sends: max-age for how long a copy stays fresh, no-cache to force a check with the server, no-store to forbid keeping it at all. It is a private cache, which makes it the one layer where personal responses may be stored.

The best candidates are JavaScript, CSS, fonts and images with a content hash in the file name. They can be cached for a year, because a new version gets a new URL. HTML should usually be revalidated on every load or cached only briefly. The catch is that you can't purge a browser: once a phone has been told to keep a file for a day, it will. Service workers add a second browser-side store, the Cache API, which is useful for offline notes but has its own update rules. Our guide to Cache-Control headers gives recipes for each type of asset.

CDN edge cache

A CDN is a shared HTTP cache spread across many cities. Students in Patna and Pune may reach different edge servers, and each edge keeps its own copies of your files, keyed by URL and a few request details. It is the natural home for static assets, video segments and public pages such as the course catalogue.

Shared means careful. Anything a CDN stores can be served to every user whose request matches, so personal pages must be marked private or bypass the CDN entirely. Unlike a browser, though, a CDN can be purged. CDN caching covers cache keys, hit ratios and Cloudflare's cache rules.

Reverse proxy and load balancer caching

In front of your application servers sits a load balancer, a reverse proxy or both. This is where teams most often assume caching that isn't there.

  • Layer 4 load balancers forward TCP or UDP connections without reading the HTTP inside them, so they cannot cache responses at all.

  • Managed layer 7 load balancers read HTTP, but that doesn't make them caches. AWS's Application Load Balancer routes by host and path, terminates TLS and runs health checks, and on AWS the cache in front of it is CloudFront's job. Google Cloud's external Application Load Balancer caches only when you switch on Cloud CDN for a backend.

  • Reverse proxies such as Nginx and Varnish do cache, and they can balance load as well. Nginx stores cached responses on disk. HAProxy also has a cache, but its documentation describes it as a minimal in-memory cache for small objects such as icons and stylesheets.

Because a reverse proxy sits right next to your application, it suits very short-lived caching of dynamic pages. Even one second of Nginx proxy_cache can turn thousands of identical requests a second into one render a second on each proxy.

Application and Redis caches

In-process memory

Inside the application server you can keep values in the process's own memory: a Python dictionary, functools.lru_cache, Django's local-memory backend or cached_property on an object. It is the fastest cache there is, with no network hop, but every process has its own copy. With 8 Gunicorn workers on each of 6 servers, a value lives in 48 separate places, and deleting it in one does nothing to the other 47. Use it for things that rarely change and can be slightly stale, such as configuration or the list of exam categories.

Shared cache: Redis or Memcached

A shared in-memory store gives every app server the same view. This is where most application caching lives: computed results such as a leaderboard or a course outline with lesson counts, rendered fragments, sessions and rate-limit counters. You decide what to store and for how long, usually with the cache-aside pattern: check the cache, and on a miss, compute the value and store it. See Redis caching for eviction and memory sizing, and Django caching for the framework side.

Database-level caching

Databases cache too, mostly without being asked. PostgreSQL keeps recently used 8 kB data pages in shared_buffers, which is 128 MB by default; its documentation suggests about 25% of RAM as a starting point on a dedicated server, and it relies on the operating system's file cache as well. A query whose pages are already in memory skips disk reads, but it still pays for parsing, planning, joining and sorting.

What PostgreSQL doesn't have is a result cache that remembers the answer to a query. For expensive, repeated queries such as leaderboards and reports, you store the result yourself, either in Redis or inside the database as a materialized view that you refresh on a schedule.

What to cache at each layer

In system design terms, these are your caching levels. The table sums up what each one is good for and how you clear it.

LayerShared byGood forKeep outHow you clear it
BrowserOne userFingerprinted JS, CSS and fonts; imagesAnything you might need to recall in a hurryYou can't; change the URL
CDN edgeEveryone near that edgeStatic files, video segments, public pages, public API responsesLogged-in pages, personal JSON, payment and OTP responsesPurge by URL, tag or prefix
Reverse proxyEveryone behind that proxyPublic dynamic pages, for one to a few secondsAnything that varies by userShort TTL, or purge where supported
In-process memoryOne worker processConfiguration, small reference listsAnything that must match across serversShort TTL or a restart
Redis or MemcachedAll app serversComputed results, fragments, sessions, countersThe only copy of important dataDelete or version the key
DatabaseAll queriesHot data pages (automatic), precomputed results in materialized viewsNothing; it is the source of truthRefresh the view

TTLs and invalidation across layers

Every cached copy needs an answer to "how long can this be wrong?" Layers make that harder, because staleness stacks. Suppose a leaderboard is cached in Redis for 60 seconds and the page showing it is cached at the CDN for five minutes. A student can see a ranking about six minutes old, because the CDN may have stored a page that was built from Redis data already a minute old. HTTP caches don't add up this way among themselves, since each passes an Age header and the next cache counts it against its own limit. The application cache, however, is invisible to HTTP.

Three habits keep this manageable:

  1. Set TTLs from the inside out. Decide how stale each kind of data may be, then keep the total across layers within that limit.

  2. Invalidate in the same order. After a change, clear the application cache first, then the reverse proxy, then the CDN. Purge the CDN first and it may refetch the old value from a layer you haven't cleared yet.

  3. Version what you can't purge. Browsers can't be purged, so static files get a new URL for each version and HTML stays short-lived.

Two failure modes deserve their own reading. Cache invalidation covers choosing TTLs and purging safely, and cache stampedes covers what happens when a popular entry expires and thousands of requests miss at the same moment.

Key takeaways

  • A request can be answered by six layers: the browser, the CDN, a reverse proxy, in-process memory, a shared cache such as Redis, and the database's own memory.

  • Only private caches such as the browser may hold personal responses. Shared caches must never serve one user's data to another.

  • Load balancers mostly don't cache: layer 4 balancers can't, and many managed layer 7 balancers don't. Reverse proxies such as Nginx and Varnish do.

  • PostgreSQL caches data pages, not query results, so store expensive results in Redis or a materialized view.

  • Staleness stacks across layers, so set TTLs and invalidate from the inside out.

Frequently asked questions

Why cache is faster than database?

A cache lookup does far less work. Redis or Memcached finds a value by its key in memory and sends it back, while a database has to parse the query, plan it, read pages, apply filters, join tables and sort the result. Databases keep hot pages in memory too, so the biggest saving usually isn't RAM versus disk. It is skipping the computation and returning an answer that was already built.

Is CDN a cache?

Yes. A CDN's core job is to act as a shared HTTP cache spread across many locations, so users get copies from a nearby edge server instead of your origin. It follows the same Cache-Control rules as other shared caches, plus its own settings. Most CDNs also do more than cache: they terminate TLS, absorb attacks, route traffic to healthy origins and can run code at the edge.

How do caches improve performance?

They improve latency, capacity and cost together. Latency drops because the copy comes from somewhere closer or cheaper to reach. Capacity rises because every hit is a request your servers and database never see, so the same hardware survives bigger spikes. Cost falls because origin bandwidth and database CPU are expensive. The size of the gain depends on the hit ratio: at 99% hits, the origin sees one request in a hundred.

Share this article

Looking for something else?

Talk to Us