Cloud Backup and Disaster Recovery: Google Cloud, AWS, Azure
Distinguish backup, replication and HA; test independent recovery access, RPO/RTO, regional loss and failback.
Content update:
Architecture and operating model
HA, replication and backup have distinct goals. HA can maintain service through component failure; replication may copy corruption; backups recover historical points. For all three clouds, define protected events: disk/zone failure, region loss, deletion, identity compromise or application corruption.
Measure RPO against the last recovered business record and RTO from the defined outage start to accepted service. A completed snapshot does not validate transactions. Managed-database backup, point-in-time and geographic options depend on engine, plan and region; do not assume universal retention or recovery time.
Separate recovery access from production identities and, where practical, attack domains. Copies, catalogues, configuration, encryption keys, licences and clean administration form one chain. Apply immutable/object-lock options according to service and mode requirements. Locking does not solve key loss or unsuitable regional design.
After failover, verify one authorised writer and client destinations. Failback includes new data and connection return. Cross-cloud recovery adds format conversion, support, egress and identity dependencies; a copy on another provider alone does not prove working DR.
- 1Protection and clean copy
- 2Independent recovery access
- 3Restore and data checks
- 4Failover / failback acceptance
Design parameters
- Protection scope
- Map assets, applications and failure types to each job/policy.
- Copy independence
- Assess region, account/project/subscription, deletion rights and key access separately.
- Measured RPO/RTO
- Record last validated transaction and first accepted operation.
- Failback data
- Plan new writes, delta transfer and old-primary reopening.
Platform implementation
Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.
| Concern | Google Cloud | Amazon Web Services (AWS) | Microsoft Azure |
|---|---|---|---|
| Backup management | Backup and DR and service-specific backup options | AWS Backup and service-specific methods | Azure Backup and service-specific methods |
| Object protection | Cloud Storage retention / Bucket Lock; verify scope. | S3 Object Lock; verify modes and retention. | Blob immutable storage; verify policy scope. |
| VM / application DR | Workload-specific replication, backup and rebuild design | AWS Elastic Disaster Recovery; verify supported sources. | Azure Site Recovery; verify supported sources. |
| Database recovery | Cloud SQL backup/PITR; engine requirements | RDS backup/PITR; engine requirements | Azure SQL / database backup options; engine requirements |
Worked example
In an exercise, outage starts at 10:00, the last recovered accepted record is 09:52 and the first successful operation is 10:47. Data loss is 8 minutes and recovery 47 minutes. With RPO 5 minutes and RTO 60 minutes, time passes but data loss fails. Do not merge both into one successful-backup label.
Prepare clean recovery targets on each candidate platform without relying on production identities. Measure DNS, keys, data and application checks; validate records before reconnection. Regional-loss and deletion exercises should cover different recovery paths.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Backup exists but recovery permission is missing | Recovery identity or keys may remain in the affected domain. | Test access and restore without primary production identity. |
| Replica is running but data is stale | Apply lag or a broken log chain may exist. | Compare last applied business records with accepted source records. |
| Data diverges after failback | Two writers or incomplete delta movement may exist. | Inspect single-writer control and last common transaction. |
Acceptance checks
- Measure RPO and RTO with separate data and timing evidence.
- Test deletion and region loss separately.
- Test keys and catalogue access without production.
- Run actual application transactions in isolated recovery.
- Verify one authorised writer after failover.
- Check new data and client targets during failback.
Related concepts
RPO and the actual loss window
RPO is the acceptable duration of data loss after an incident. A backup schedule alone does not prove it: failed jobs or delayed replication to another site can enlarge the window. Compare the incident time with the latest usable, consistent recovery point. If an incident occurs at 14:00 and the verified copy is from 13:20, the observed loss window is 40 minutes. Set targets per application; a file archive and a database receiving continuous orders may have different requirements.
RTO and end-to-end recovery time
RTO specifies how soon a service must become usable after an incident. Downloading a backup or booting a VM accounts for only part of that duration. Record detection, approval, infrastructure preparation, data transfer, application startup and business validation separately. Define exactly what starts and stops the test clock. The same technical restore duration can produce different business interruptions when dependencies or access approvals introduce additional waiting.
Immutability and deletion authority
Immutable retention restricts modification or deletion for a defined period. An application promise, storage-enforced protection and physical isolation are different controls. Review lock duration, clock handling, administrative privileges and deletion paths together. Whether data is automatically deleted after expiry depends on the product policy. On a controlled test copy, attempt deletion using ordinary and privileged accounts, then verify recovery from the protected copy. Protection against deletion does not replace a readability or restore test.
Replication scope
Replication transfers data or state changes to another copy. Synchronous transfer can introduce latency and connectivity dependence; asynchronous transfer can lag behind. A current replica is not necessarily protected from incorrect changes: deletion and corruption can propagate too. Compare lag with the application’s last committed operation, not merely connection status. Define which copy receives write authority after the source is lost and how the old node is reconciled when it returns. Report replication lag and the actual recoverable point separately.
Dependencies and restart order
Services commonly depend on identity, DNS, time, networking, databases and licensing. Record a dependency graph describing conditions for operation, not merely an equipment list. Recovery order follows that graph; circular dependencies may require emergency access paths. Distinguish restored infrastructure from resumed business activity. Assign an owner, validation method and alternative access path to each dependency. Test assumptions by deliberately making one component unavailable in a controlled end-to-end exercise.
Application consistency
A copy that boots does not prove application-data consistency. Operating-system caches, database logs and write ordering across disks or services affect the result. A crash-consistent copy resembles recovery after an unexpected shutdown; an application-consistent copy follows supported application preparation and write coordination. Validate transaction integrity, relationships between records and application behaviour after recovery, rather than only counting files. Confirm backup integration, application version and reported errors before assuming that consistency was achieved.
Primary documentation
How this guide was prepared
Doz Teknoloji Technical Team. This guide uses provider documentation, shared architecture principles and examples with stated assumptions. Calculations and scenarios illustrate the method; they do not claim completed customer tests or results. Verify versions, regions, service scope and support conditions for implementation.
Related cloud guides
Google Cloud, AWS and Azure: Choosing a Cloud Platform
Evaluate Google Cloud, AWS and Azure through workload, service model, data location, security, recovery and total cost.
Cloud Identity and Access: Google Cloud, AWS, Azure
A three-platform guide to human and workload identities, temporary access, least privilege, resource boundaries and access validation.
Cloud Networking: Google Cloud VPC, AWS VPC and Azure VNet
Design and troubleshoot addressing, platform boundaries, routing, private access, firewall controls and hybrid connectivity.
Cloud Migration: Google Cloud, AWS and Azure Planning
Plan a controlled cloud migration through discovery, dependencies, service selection, synchronisation, cutover and rollback.
Cloud Cost and FinOps: Google Cloud, AWS, Azure
Manage compute, storage, egress, logging and licensing alongside budgets, rightsizing and commitments.
Cloud deployment, operation and support services
We assess platform choice alongside your existing products, workloads and operational needs.