RTO and RPO: planning disaster recovery that works
Set recovery time and recovery point objectives per system, choose between backup and restore, pilot light, warm standby and active/active, apply 3-2-1 backups and run drills.
On this page 8 sections
RTO, the recovery time objective, is the longest a system may stay down after a disaster; RPO, the recovery point objective, is how much recent data you can afford to lose, measured in time. A four-hour RTO with a 15-minute RPO means: back within four hours, losing at most the last 15 minutes of changes. Set both per system rather than for the whole platform, choose the cheapest recovery strategy that meets each pair, and prove the numbers with regular drills.
RTO and RPO in plain words
AWS's disaster-recovery whitepaper defines RTO as the maximum acceptable delay between the interruption of service and its restoration, and RPO as the maximum acceptable time since the last data recovery point. Picture a timeline:
- Your last good copy of the data was taken at 9:45.
- Disaster strikes at 10:00. Everything written between 9:45 and 10:00 is at risk: that gap is what your RPO has to cover.
- Service is restored at 11:30. The 90 minutes of downtime is what your RTO has to cover.
Two points are easy to miss. First, both are objectives, decided by the business in advance; what actually happens in an incident may be better or worse. Second, the disaster isn't always an outage. If a bad deploy corrupts data at 10:00 and nobody notices until 2:00, your recovery point must be before 10:00, and replicas won't help, because they faithfully copied the corruption. Only backups taken before the damage can.
Setting targets per system
A learning platform isn't one system with one tolerance for loss. Losing a payment record is serious; losing yesterday's analytics rollup is an inconvenience. The table below is illustrative; your figures should come from asking what an hour of downtime, or an hour of lost data, costs in refunds, retests, support load and reputation:
| System | What failure means | Illustrative RTO | Illustrative RPO | Likely approach |
|---|---|---|---|---|
| Payments and enrolments | Money taken but no access given; reconciliation work | 30 minutes | Near zero | Replicated database with point-in-time recovery; reconcile against the payment gateway's records |
| Test attempts and answers | Students lose answers and must resit | 15 minutes during exam windows | Under a minute | Replicated database; answers also cached on the device until acknowledged |
| Recorded lectures | Students can't study | 1 hour | 24 hours, since lectures rarely change | Replicated object storage, originals kept, CDN in front |
| Live classes | A class can't start | 15 minutes, or reschedule | Not applicable | Backup ingest and a second delivery path |
| Analytics and reports | Dashboards arrive late | 1 to 2 days | 24 hours | Nightly backups |
| Caches and search indexes | Slower pages until rebuilt | Hours | None needed | Rebuild from the source of truth |
Targets can also depend on the calendar. The whitepaper gives the example of a payroll system that matters most just before payday; for a coaching platform, exam weeks and result days are the equivalent, and it makes sense to set stricter targets for those windows. It also notes that if a recovery strategy costs more than the loss it prevents, you shouldn't build it unless something else, such as regulation, requires it.
Four disaster recovery strategies
The AWS whitepaper describes four strategies, in rising order of cost and complexity. The same ideas apply on any cloud:
| Strategy | What runs in the recovery site | Typical recovery | Cost |
|---|---|---|---|
| Backup and restore | Nothing until needed; data and infrastructure are restored from backups and code | Hours | Lowest |
| Pilot light | Data is replicated continuously and core services such as databases are on; application servers are switched off | Tens of minutes | Low |
| Warm standby | A scaled-down but fully working copy that can take traffic immediately at reduced capacity | Minutes | Medium |
| Multi-site active/active | Full production in two or more regions, all serving users | Near zero for most failures | Highest |
The difference between pilot light and warm standby trips people up. A pilot light can't serve requests until you switch on and scale up the application tier; a warm standby is already serving, just small. Active/active adds the hardest problem of all, handling writes in more than one region, and even it needs backups, because corruption and accidental deletes replicate everywhere instantly.
Backups: the 3-2-1 rule
The 3-2-1 rule, as CISA's guidance states it, is to keep three copies of any important file (one primary and two backups), on two different media types, with one copy offsite. On a cloud platform that translates to:
- The primary database, plus automated backups with point-in-time recovery in the same region, which cover most day-to-day accidents.
- Copies of those backups in a second region, so a regional failure doesn't take out both.
- At least one copy in a separate account with restricted access, stored immutably. Amazon S3 Object Lock, for example, stores objects write-once-read-many, and in compliance mode not even the account's root user can delete a locked version before its retention date. That protects backups from ransomware and from a compromised admin account.
For a Postgres database, our guide to Postgres point-in-time recovery covers base backups, WAL archiving and restoring to the second before a mistake.
Region failures and DNS
Most incidents are smaller than a region going down, and running across several availability zones within one region already covers a data-centre failure. Region-level recovery needs a second region with your data in it. For Indian platforms that want every copy in India, AWS, for example, has two Indian regions, Mumbai and Hyderabad.
Failing over between regions involves three moves: promote the database in the recovery region, scale up the application tier, and send traffic there. For the database, read replicas in the recovery region can be promoted to primary; for non-Aurora RDS instances, AWS notes that promotion takes a few minutes and includes a reboot. For traffic, DNS failover with health checks is the common tool: keep the TTL on those records short, 60 or 120 seconds being a common choice, and rely on the provider's data plane, such as health checks and routing controls, rather than console changes that depend on a control plane which may itself be struggling.
Automatic failover sounds attractive, but a false alarm costs you downtime and possibly data. The whitepaper suggests that many teams start failover manually, with every step after that decision automated, so that the switch itself is a single action.
Testing your plan
AWS's guidance puts it bluntly: the only recovery path that works is the one you test frequently. A practical schedule:
- Restore tests, monthly. Restore the latest backup into a scratch environment automatically, time it, check the newest record in it, and run a smoke test. The duration is your real RTO for backup and restore; the age of the newest record is your real RPO.
- Failover game days, quarterly. In a quiet week, never exam season, promote a replica or fail over a region, and record every timing.
- Tabletop exercises. Walk through scenarios such as "the CDN fails during a live class" or "someone ran a DELETE without a WHERE clause", and decide who does what and who tells students.
- Drift checks. Make sure the recovery region has current images, configuration and enough service quota to scale up.
Record each drill in a simple scorecard. An illustrative one:
| Scenario | Target RTO | Measured | Target RPO | Measured | Action |
|---|---|---|---|---|---|
| Rebuild the main database from backups (the worst case) | 2 h | 3 h 10 min | 5 min | 4 min | Take incremental backups more often to shorten log replay |
| Promote attempts replica in the recovery region | 15 min | 11 min | 1 min | 20 s | None |
| Switch video delivery to the second CDN | 15 min | 35 min | n/a | n/a | Tokens failed on the backup CDN; fix and re-test |
A drill that misses its target isn't a failure; it's the cheapest way to find out. Tie the results back to your alerting, too: our guide to monitoring vs observability covers the signals that tell you a disaster has started, and multi-CDN failover covers the video delivery path.
Key takeaways
- RTO is how long you can be down; RPO is how much recent data you can lose. Set both per system.
- Pick the cheapest of backup and restore, pilot light, warm standby or active/active that meets each target.
- Replication isn't backup: keep 3-2-1 copies, including one immutable copy in a separate account.
- Keep failover DNS TTLs short, use data-plane controls, and automate every step after the decision to fail over.
- Test restores and failovers on a schedule, measure real RTO and RPO, and fix the gaps.
Frequently asked questions
What is RTO and RPO in backup?
In backup planning, RPO is set mainly by how often you back up: nightly dumps mean up to a day of lost data, while continuous log archiving can bring it down to minutes or seconds. RTO is set by how fast you can restore, which depends on data size, storage speed and how much log must be replayed. Measure both with real restore tests, not estimates.
What is RTO and RPO in AWS?
AWS uses the standard definitions in its Well-Architected guidance and disaster-recovery whitepaper, and maps them to four strategies: backup and restore, pilot light, warm standby and multi-site active/active. Services then set practical limits. For example, Amazon RDS uploads database transaction logs to S3 every five minutes, which bounds how recent a point-in-time restore can be.
What is RTO and RPO in DR?
In a disaster recovery plan, RTO and RPO are the targets each system's recovery must meet, taken from a business impact analysis. They decide which DR strategy each system gets, how often backups run, whether data is replicated to another region, and what the drills must prove. A DR plan without explicit RTO and RPO values can't be tested properly.
What is the difference between RTO and RPO?
RTO measures time without service: how long after a disaster the system must be working again. RPO measures data loss: how far back in time your recovered data may be. They are independent. A read-only lecture library can accept a long RPO but needs a short RTO, while a payments ledger needs a near-zero RPO even if a slightly longer outage is tolerable.