Multi-CDN and CDN failover for video platforms
Why one CDN is a single point of failure on exam and launch days, how DNS, player-side and content-steering failover compare, and what must match across CDNs.
On this page 8 sections
Multi-CDN means delivering the same videos through two or more content delivery networks, with a way to move viewers between them when one slows down or fails. For a video platform, a single CDN is a single point of failure: if it has a bad hour during a results-day live stream or an all-India mock test, so do you. Traffic can be switched in DNS, in the player, or by a steering service, and every option depends on the CDNs sharing the same URLs, tokens and cache rules.
How a CDN delivers video
A CDN runs servers in many locations, called edges or points of presence. A student's request is routed to a nearby edge, either by DNS answering with a nearby server's address or by anycast, where many locations share one IP address. The edge looks up the request in its cache using a cache key, usually built from the URL:
- On a hit, it serves the segment from memory or disk.
- On a miss, it asks a parent cache, often a regional tier, which may in turn ask an origin shield. Only if every tier misses does the request reach your origin, typically object storage or a packager.
Recorded video suits this model well. A segment of a published lecture never changes, so it can be cached for a long time and served to thousands of students from one fetch. Our guide to CDN caching covers cache keys and hit ratios in detail.
Why one CDN isn't enough
CDNs are very reliable overall, but a large, distributed system fails in partial and local ways that still hurt:
- a problem between the CDN and one internet provider degrades video for that ISP's customers in one region, while everyone else is fine;
- a configuration or certificate change at the CDN goes wrong;
- a regional outage takes out the locations that serve a large share of your students;
- a huge live event pushes capacity in some locations, and performance drops just when everyone is watching.
For a coaching institute, the worst hour is predictable: the live class before an exam, the answer-key stream, the free mock test you advertised for a month. A second CDN turns a provider's bad hour into a routing decision rather than an outage. It also gives you commercial leverage when contracts come up for renewal.
DNS-based vs client-side switching
| Method | How it switches | Speed | Granularity | Effort |
|---|---|---|---|---|
| DNS steering | Your DNS provider answers with CDN A or CDN B, by weight, latency or health check | Minutes: resolvers cache answers for the record's TTL, and a playing video keeps its current CDN until it looks the name up again | Per resolver, which often means per ISP or region | Low; no player changes |
| Player-side failover | The player knows several base URLs and retries a failed or slow segment on the next one | Seconds, within a single session | Per viewer, per request | Moderate; player logic and testing |
| Content steering | A steering server tells players which CDN to prefer; players re-check it periodically | Seconds to minutes, under central control | Per session, by any rule you like | Higher; a steering service plus player support |
With DNS steering, keep the TTL short on the records you switch. Amazon's Route 53 guidance calls 60 or 120 seconds a common choice for health-checked failover records. Even then, DNS only affects new lookups; students halfway through a lecture stay where they are until the player resolves the name again.
Player-side failover fixes that. Apple's HLS authoring specification says streams should support failover, for example by listing duplicate variants in the multivariant playlist. Content steering, defined in the HLS second-edition draft and in an equivalent DASH-IF specification, adds central control. The multivariant playlist names a steering server and tags each variant with a pathway, usually one per CDN (other attributes trimmed):
#EXTM3U
#EXT-X-CONTENT-STEERING:SERVER-URI="https://steer.example.com/v1?session=s123",PATHWAY-ID="cdn-a"
#EXT-X-STREAM-INF:BANDWIDTH=620000,RESOLUTION=640x360,PATHWAY-ID="cdn-a"
https://a.cdn.example.net/lectures/42/360p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=620000,RESOLUTION=640x360,PATHWAY-ID="cdn-b"
https://b.cdn.example.org/lectures/42/360p/index.m3u8
The steering server returns a small JSON document with a PATHWAY-PRIORITY list and a time-to-live; the player uses the first pathway on the list that appears in its playlist, and fetches the document again when the time-to-live runs out. Change the order on the server, and players move over within one refresh. Many teams combine methods: DNS or steering for the default split, and player retries as the safety net.
Consistent tokens, keys and cache rules
Most multi-CDN failures are self-inflicted: the backup CDN works in testing but rejects real students because something differs. Check these on every CDN:
- Identical paths. The same path should return the same object from either CDN, and your origin shouldn't care which CDN is asking.
- Access tokens. Each CDN has its own signed-URL or token format, so your playback API must be able to issue a valid one for every pathway. The Common Access Token, standardised by the Consumer Technology Association as CTA-5007, aims to give CDNs one token format. Our guide to signed URLs covers expiry and binding.
- Cache keys without tokens. Strip per-student tokens and session parameters from the cache key on every CDN. Otherwise each student gets a unique cache key, every request becomes a miss, and a failover can flood your origin.
- Matching cache lifetimes. Long for segments, short for live playlists, and the same on both CDNs.
- TLS and CORS. A valid certificate for your custom hostname on each CDN, and the same cross-origin headers, or web players will fail on the backup.
- Purging. A replaced lecture must be purged on every CDN, not just the primary.
- Origin protection. Each CDN sends its own cache misses to your origin, so two CDNs mean more origin traffic. AWS's documentation on CloudFront Origin Shield describes one answer: put a shared caching layer between all CDNs and the origin, so a segment is fetched from storage once.
Monitoring CDN health from real users
A CDN's own status page tells you about large incidents. What you need is the view from your students' players, broken down by CDN, internet provider and city: segment download times, error rates, measured throughput, startup time and rebuffering. Common Media Client Data (CMCD, CTA-5004) helps here, because players attach values such as buffer length, measured throughput and a session ID to each segment request, and they show up in the CDN's logs. Synthetic probes from the cities and ISPs your students use fill the gaps at quiet hours.
Turn the data into explicit switching rules, and add hysteresis so traffic doesn't bounce between CDNs every minute. An illustrative rule set, per ISP and city over five-minute windows:
| Signal | Illustrative threshold | Action |
|---|---|---|
| Segment error rate on CDN A | Above 2%, and at least three times CDN B's rate | Steer that ISP and city to CDN B |
| 90th-percentile segment download time on CDN A | More than half the segment duration | Shift weight gradually towards CDN B |
| Rebuffering ratio across both CDNs | Above 1% for 10 minutes | Page the on-call engineer: the problem may be the origin or the encoder, not a CDN |
| CDN A healthy again | Normal for 30 minutes | Move traffic back in steps, not all at once |
Tune the numbers to your own traffic; the point is to decide them calmly before the incident, not during it.
Cost and contracts
- Volume commitments split. CDN pricing usually improves with committed volume, and splitting traffic lowers your volume with each vendor. Negotiate with the split in mind.
- Keep the backup warm. A CDN that carries no traffic has cold caches and untested configuration. Send it a steady share of real traffic so you know it works and its caches hold your popular lectures.
- Origin and shield costs rise with each CDN that fetches from you, unless a shared shield absorbs the misses.
- Engineering time goes into token issuance, player logic, monitoring and regular failover drills.
For many platforms, a primary CDN plus a warm secondary, with player-side failover and a tested manual switch, delivers most of the resilience at a fraction of the complexity of fully automatic, performance-based steering. Our guide to scaling a video learning platform shows where the CDN fits among the other layers, and RTO and RPO planning covers the wider recovery picture.
Key takeaways
- A single CDN is a single point of failure, and the hours it fails are the ones everyone is watching.
- DNS steering is simple but slow and coarse; player-side failover acts per viewer in seconds; content steering adds central control.
- Make every CDN interchangeable: same paths, valid tokens for each, tokens stripped from cache keys, same cache lifetimes and certificates.
- Judge CDN health from real players, by ISP and city, and decide switching rules before you need them.
Frequently asked questions
What is multi CDN?
Multi-CDN is the practice of delivering content through two or more CDN providers at once, with rules for splitting traffic between them and moving it when one performs badly. It protects against a single provider's outage or regional trouble, lets you route each region or ISP to whichever CDN performs best there, and strengthens your position in price negotiations.
How do CDNs work?
A CDN places caching servers in many locations and routes each user to a nearby one, using DNS or anycast. If that server already holds the requested file, it serves it directly; if not, it fetches it from a parent cache or from your origin, stores a copy, and serves later requests from that copy. The result is shorter distances for users and far less load on your servers.
How do CDN servers work?
An individual CDN edge server terminates the user's TLS connection, builds a cache key from the request, and looks it up in memory and on local disks. On a hit it responds immediately; on a miss it fetches the object from an upstream cache or the origin, stores it and responds. It also enforces rules such as token checks and cache lifetimes, and evicts rarely used objects when space runs low.
Is DNS failover enough for video?
Usually not on its own. DNS changes reach users only as resolvers' cached answers expire, and a student already watching keeps fetching segments from the same CDN until the player looks the hostname up again. Pair DNS steering with player-side failover, so a failing segment request is retried on another CDN within seconds, and test the switch regularly.