Skip to main content
CLOUD ARCHITECTURE

Google Cloud, AWS and Azure: Choosing a Cloud Platform

Evaluate Google Cloud, AWS and Azure through workload, service model, data location, security, recovery and total cost.

Content update:

Architecture and operating model

Platform selection is requirements matching rather than brand ranking. Start with transactions, data volume, latency tolerance, existing licences and operating skills. Then test whether virtual machines, containers, serverless or managed services fit. A suitable service on one provider may have different names and defaults on another.

With IaaS virtual machines, OS and application maintenance largely remain customer responsibilities; managed services shift some infrastructure tasks to the provider. Identity, data permissions, configuration and application responsibilities still need review. Managed services do not automatically complete backup, access policy or cost governance.

Region, zone, service availability and quotas are separate inputs. Equal vCPU counts do not imply equal CPU generations or performance. Keeping primary data in one region does not prove logs, backups and management data remain there. Review data paths per service.

Multiple clouds add identity, networking, observability, data-transfer and response costs alongside potential independence. If the workload does not justify them, a well-designed architecture on one platform with tested export paths may be more practical. Exit planning must be measurable in either case.

  1. 1Requirements
  2. 2Service model
  3. 3Pilot and acceptance
  4. 4Operation and exit
Compare all three platforms against the same acceptance criteria.

Design parameters

Service model
Document owners for OS, runtime, data and patches per service.
Performance target
Compare p95/p99, errors and usage under equivalent data and concurrency.
Data and recovery
Evaluate location, keys, backup and RPO/RTO together.
Exit and cost
Include licensing, egress, conversion and operations rather than VM-hours alone.

Platform implementation

Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.

ConcernGoogle CloudAmazon Web Services (AWS)Microsoft Azure
Virtual machineCompute EngineAmazon EC2Azure Virtual Machines
Object storageCloud StorageAmazon S3Azure Blob Storage
Managed relational dataCloud SQL; verify supported engines/versions.Amazon RDS; review engine choices and behaviour.Azure SQL and Azure Database services; select against engine needs.
Container operationGoogle Kubernetes Engine (GKE)Amazon EKS; alternative ECS has different scope.Azure Kubernetes Service (AKS)

Worked example

Assume an ERP with 2 TB of data, 150 concurrent users and a 300 ms p95 target for critical queries. Build pilots representing the same transaction set, distribution and network distance as closely as possible. Start with comparable CPU, RAM and storage, then right-size from measured needs on each platform.

Include production, staging, backups, logs, licences, support and egress in the estimate. Do not declare a cheaper option the winner if it fails technical acceptance. Test recovery, denied access and export alongside benchmarks.

Troubleshooting

ObservationLikely cause / distinctionVerification
Equal sizes produce different performanceCPU generation, storage limits or network paths may differ.Compare query plans, storage latency and instance characteristics under equal load.
Service is absent in the selected regionProduct and region coverage may differ.Validate availability, quotas and data-path requirements before design.
Pilot is cheap but production is expensiveHA, backup, logs or egress may be omitted from the pilot.Map the complete service scope and actual monthly usage units.

Acceptance checks

  1. Compare all three platforms against the same acceptance criteria.
  2. Document each service’s responsibility boundary.
  3. Measure latency and errors under realistic load.
  4. Test backup and recovery on a separate target.
  5. Record total cost and egress assumptions.
  6. Test export paths for data, configuration and keys.

Related concepts

Capacity and usable headroom

Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.

Latency distribution

Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.

Dependencies and restart order

Services commonly depend on identity, DNS, time, networking, databases and licensing. Record a dependency graph describing conditions for operation, not merely an equipment list. Recovery order follows that graph; circular dependencies may require emergency access paths. Distinguish restored infrastructure from resumed business activity. Assign an owner, validation method and alternative access path to each dependency. Test assumptions by deliberately making one component unavailable in a controlled end-to-end exercise.

Availability versus recovery

High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.

Cost boundaries

Cost includes more than purchase price or a monthly resource bill. Licensing scope, storage, egress, backup, support, operating effort and disruption impact are separate items. Compare equal service levels, capacity and time periods; a cheaper option may exclude an obligation. Compare forecast and actual spending by resource or business unit. Assess underuse risk for commitments and uncontrolled growth for flexible models. Verify pricing and licensing conditions against official documents at decision time.

Version and support lifecycle

Installability does not prove production support. Review the compatibility chain across OS, application, drivers, extensions and management tools. Update plans should record version, support end, restart needs and rollback methods. An unrepresentative test environment can produce misleading results. Validate service health and existing workflows after a change, not just version numbers. Remember that pinning a version can also prevent future security fixes.

Primary documentation

How this guide was prepared

Doz Teknoloji Technical Team. This guide uses provider documentation, shared architecture principles and examples with stated assumptions. Calculations and scenarios illustrate the method; they do not claim completed customer tests or results. Verify versions, regions, service scope and support conditions for implementation.

Enterprise IT Product Sales, Licensing and Deployment
Enterprise IT Project & Solution Scenarios
View all related content
Text on WhatsApp
Copied!