Cloud Networking: Google Cloud VPC, AWS VPC and Azure VNet
Design and troubleshoot addressing, platform boundaries, routing, private access, firewall controls and hybrid connectivity.
Content update:
Architecture and operating model
Cloud networks are not direct copies of physical VLANs. Google Cloud VPC is global with regional subnets. AWS VPC is regional with subnets in one Availability Zone. Azure VNet is regional with subnets within that VNet. These differences affect placement, addressing and multi-region design; do not copy one subnet diagram blindly across platforms.
Routes provide paths and firewalls grant access; neither replaces the other. AWS security groups are stateful and network ACLs stateless. Google Cloud VPC firewall rules and Azure NSGs are stateful, with differing targets, priorities and defaults. Validate return paths, NAT, ports and application listeners together.
Private endpoints or API access involve DNS, source networks, permissions and closure of public access, not just private IPs. Peering does not automatically create all transit paths. Hybrid VPN/dedicated connectivity must address overlapping CIDRs, BGP, MTU and failover.
Measure symmetry, scaling and failure effects in central inspection designs. NAT is not a firewall governing every inbound path. Separate management and application traffic; correlate flow logs with identity and application records.
- 1CIDR and scope
- 2DNS and route
- 3Firewall and private service
- 4Application transaction
Design parameters
- Addressing and growth
- Allocate non-overlapping on-premises, cloud, container and DR CIDRs.
- Routes and return path
- Record effective routes, NAT and source-address behaviour in both directions.
- DNS and private access
- Verify resolvers, zones and public/private answers from actual clients.
- Failure acceptance
- Measure existing sessions and new transactions separately during path loss.
Platform implementation
Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.
| Concern | Google Cloud | Amazon Web Services (AWS) | Microsoft Azure |
|---|---|---|---|
| Network and subnet | Global VPC / regional subnet | Regional VPC / zonal subnet | Regional VNet / VNet subnet |
| Traffic control | VPC firewall rules and firewall-policy options | Security groups; network ACLs where required | Network Security Groups; Azure Firewall where needed |
| Hybrid connection | Cloud VPN / Cloud Interconnect | Site-to-Site VPN / Direct Connect | VPN Gateway / ExpressRoute |
| Private service access | Private Service Connect and service-specific access | VPC endpoints / AWS PrivateLink; endpoint type matters. | Private Endpoint / Azure Private Link |
Worked example
With an on-premises 10.20.0.0/16 network, choose non-overlapping ranges for three candidate clouds. Restrict the application to a private database and terminate web ingress at a load balancer. Use a separate management path. Validate DNS, routes and TCP/TLS from actual clients.
If the private database name resolves publicly, adding a firewall rule alone does not explain the problem. Check private DNS visibility and resolver forwarding. Record tunnel loss and the application’s first successful transaction on synchronised clocks.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| DNS is correct but TCP fails | Route, return path, firewall or listener may be wrong. | Compare effective routes and both-end flow records with listeners. |
| Private endpoint resolves publicly | The client may lack private-zone visibility. | Compare authoritative DNS and client resolver answers for the same name. |
| Small packets work but large transfers stall | MTU or path-MTU discovery may be involved. | Measure packet sizes, retransmissions and ICMP behaviour under controlled conditions. |
Acceptance checks
- Validate access across DNS, routing, security and application layers.
- Check CIDR overlap on normal and DR paths.
- Test allowed and denied source/target pairs.
- Resolve private service names from actual clients.
- Measure new and existing sessions during path loss.
- Verify management access and flow logging.
Related concepts
Layer 2 and Layer 3 boundaries
A VLAN creates a separate broadcast domain; routing and access policies govern communication between VLANs. Review trunk allowed lists, access-port assignment and gateway placement together. A network diagram should show where packets are routed and filtered, not merely which cables connect devices. Sharing a switch does not require sharing privileges. When access fails, confirm VLAN and addressing first, then gateway, route and policy matching.
DNS caching and dependencies
DNS results depend on client and recursive-resolver caches as well as authoritative data. Changing a TTL does not retroactively shorten already cached answers. A/AAAA, CNAME, MX and TXT records serve different purposes. Test internal and external views separately; split DNS may intentionally return different results. Record resolver, response type, TTL and query time when troubleshooting. Successful access by IP address alone does not prove correct DNS configuration.
MTU along the path
MTU concerns packet size on a link; tunnelling and encapsulation headers can reduce space available for useful payload. Small requests working while large transfers stall may suggest an MTU issue, but this is not proof. TCP MSS and interface MTU are distinct. Compare both endpoints and tunnel interfaces. Identify the affected segment with controlled tests instead of changing every device, and remember that blocked ICMP error messages can complicate path-MTU discovery.
Trust boundary and failure domain
A trust boundary separates components governed by different access decisions; a failure domain groups resources that one event can affect together. Two VLANs do not create a strong trust boundary when routing between them is unrestricted. Backups in different folders still share a failure domain if one administrator can delete both. Assess physical location, identity provider, management account, network path and power source separately. Test which access paths and recovery options remain available when a component is lost.
Latency distribution
Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.
Availability versus recovery
High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.
Primary documentation
How this guide was prepared
Doz Teknoloji Technical Team. This guide uses provider documentation, shared architecture principles and examples with stated assumptions. Calculations and scenarios illustrate the method; they do not claim completed customer tests or results. Verify versions, regions, service scope and support conditions for implementation.
Related cloud guides
Google Cloud, AWS and Azure: Choosing a Cloud Platform
Evaluate Google Cloud, AWS and Azure through workload, service model, data location, security, recovery and total cost.
Cloud Identity and Access: Google Cloud, AWS, Azure
A three-platform guide to human and workload identities, temporary access, least privilege, resource boundaries and access validation.
Cloud Migration: Google Cloud, AWS and Azure Planning
Plan a controlled cloud migration through discovery, dependencies, service selection, synchronisation, cutover and rollback.
Cloud Cost and FinOps: Google Cloud, AWS, Azure
Manage compute, storage, egress, logging and licensing alongside budgets, rightsizing and commitments.
Cloud Backup and Disaster Recovery: Google Cloud, AWS, Azure
Distinguish backup, replication and HA; test independent recovery access, RPO/RTO, regional loss and failback.
Cloud deployment, operation and support services
We assess platform choice alongside your existing products, workloads and operational needs.