Cloud Cost and FinOps: Google Cloud, AWS, Azure
Manage compute, storage, egress, logging and licensing alongside budgets, rightsizing and commitments.
Content update:
Architecture and operating model
FinOps links costs to business outcomes. Track cost per transaction, utilised capacity and resource ownership alongside monthly totals. Labels/tags allocate costs to teams and workloads; provider coverage and billing delays vary. Define allocation rules for untagged resources and shared infrastructure.
Resizing a VM may not reduce the total bill proportionally. Measure disks, snapshots, managed databases, logs, NAT, load balancers and egress separately. A stopped VM may retain billable disks or addresses depending on service behaviour. Check dependencies and retention before deleting resources.
Right-size against performance targets and failure reserves. Commitments, reservations and savings options differ in duration, coverage and flexibility. They are not identical products and savings are not guaranteed for every workload. Do not commit variable test demand as if it were stable production demand.
Budget alerts and stopping usage are different controls. Azure budgets and AWS notifications require separately designed actions. Google Cloud alerts-only budgets do not stop spending; separate spend caps for eligible services require scope and reporting-delay review. Shutdown mechanisms can disrupt production and need ownership, exceptions and restart planning.
- 1Usage and ownership
- 2Cost allocation
- 3Measured optimisation
- 4Budget and review
Design parameters
- Usage units
- Map VM-hours, GB-months, requests, I/O and egress to current billing units.
- Performance reserve
- Right-size while meeting p95/p99 and N-1 needs; average CPU alone is insufficient.
- Commitment scope
- Document duration, service/region coverage and unused-commitment risk.
- Alert and action
- Test budget scope, delay, recipients and action permissions separately.
Platform implementation
Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.
| Concern | Google Cloud | Amazon Web Services (AWS) | Microsoft Azure |
|---|---|---|---|
| Visibility | Cloud Billing reports / export | AWS Cost Explorer / billing exports | Microsoft Cost Management |
| Budget | Cloud Billing budgets; alerts-only and eligible spend caps differ. | AWS Budgets; verify notifications and budget-action scope. | Cost Management budgets; action automation is configured separately. |
| Commitment | Committed use discounts; terms differ by service. | Savings Plans / Reserved Instances; coverage differs. | Savings plan / reservations; coverage differs. |
| Allocation | Labels and billing scope | Cost allocation tags and account boundary | Tags and subscription/resource scope |
Worked example
Use a hypothetical unit cost of 1 per hour, not a current price. Continuous operation for 30 days is 720 hours; 22 days × 10 hours of testing is 220 hours. Compute differs by 500 units, while storage, backups, logs and commitments may remain. Estimate all three providers with the same schedule, service scope and region.
A total cost of 200 units for 10,000 transactions is 0.02 per transaction. If next month costs 250 for 20,000 transactions, unit cost is 0.0125. Efficiency can improve while total spending rises. Do not interpret outcomes from the bill total or one resource percentage alone.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| VM stops but costs continue | Disks, addresses, snapshots or commitments may remain billable. | Compare actual meters/resources against usage schedules. |
| Budget is exceeded but resources run | The budget may be alerts-only or actions may have different scope. | Inspect budget type, reporting delay and action results. |
| Discount exists but savings are absent | Commitments may be unused or costs may lie in another service. | Measure coverage, utilisation and on-demand cost together. |
Acceptance checks
- Validate cost reduction while preserving service acceptance criteria.
- Audit resource owners and tag coverage.
- Calculate non-compute cost items.
- Measure actual baseline usage before commitments.
- Validate alerts and actions separately in controlled tests.
- Report total and per-transaction costs together.
Related concepts
Cost boundaries
Cost includes more than purchase price or a monthly resource bill. Licensing scope, storage, egress, backup, support, operating effort and disruption impact are separate items. Compare equal service levels, capacity and time periods; a cheaper option may exclude an obligation. Compare forecast and actual spending by resource or business unit. Assess underuse risk for commitments and uncontrolled growth for flexible models. Verify pricing and licensing conditions against official documents at decision time.
Capacity and usable headroom
Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
Availability versus recovery
High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.
Policy lifecycle
A policy needs management as its scope changes, not just when it is enabled. Separate drafting, observation, pilot deployment, enforcement and periodic review. Every exception needs an owner, reason, scope and expiry date. A successful pilot can still cause production false positives if test data does not represent real user behaviour. Identify affected users and applications before deploying with measurable acceptance criteria. A rollback path should identify who may reverse the change and under which conditions, rather than merely locating an interface button.
Version and support lifecycle
Installability does not prove production support. Review the compatibility chain across OS, application, drivers, extensions and management tools. Update plans should record version, support end, restart needs and rollback methods. An unrepresentative test environment can produce misleading results. Validate service health and existing workflows after a change, not just version numbers. Remember that pinning a version can also prevent future security fixes.
Primary documentation
How this guide was prepared
Doz Teknoloji Technical Team. This guide uses provider documentation, shared architecture principles and examples with stated assumptions. Calculations and scenarios illustrate the method; they do not claim completed customer tests or results. Verify versions, regions, service scope and support conditions for implementation.
Related cloud guides
Google Cloud, AWS and Azure: Choosing a Cloud Platform
Evaluate Google Cloud, AWS and Azure through workload, service model, data location, security, recovery and total cost.
Cloud Identity and Access: Google Cloud, AWS, Azure
A three-platform guide to human and workload identities, temporary access, least privilege, resource boundaries and access validation.
Cloud Networking: Google Cloud VPC, AWS VPC and Azure VNet
Design and troubleshoot addressing, platform boundaries, routing, private access, firewall controls and hybrid connectivity.
Cloud Migration: Google Cloud, AWS and Azure Planning
Plan a controlled cloud migration through discovery, dependencies, service selection, synchronisation, cutover and rollback.
Cloud Backup and Disaster Recovery: Google Cloud, AWS, Azure
Distinguish backup, replication and HA; test independent recovery access, RPO/RTO, regional loss and failback.
Cloud deployment, operation and support services
We assess platform choice alongside your existing products, workloads and operational needs.