Skip to main content

Troubleshooting · Servers and Virtualization

Diagnose VM Memory Ballooning and Host Swap

A balloon driver can reclaim guest memory for the host. Guest pressure and host swapping are different layers. Host free memory, guest process RSS or guest available memory alone cannot…

Technical review:

Architecture and operating model

A balloon driver can reclaim guest memory for the host. Guest pressure and host swapping are different layers. Host free memory, guest process RSS or guest available memory alone cannot explain both.

Collect host PSI, swap-in/out, guest statistics and application latency on one timeline. Balloon statistics depend on guest-driver and polling support; missing/zero fields do not necessarily mean no usage. Check hugepages and pinned memory separately.

  1. 1Guest demand
  2. 2Balloon/reclaim
  3. 3Host memory pressure
  4. 4Application latency
Align host/guest measurements.

Design parameters

Pressure measurement
Record swap activity on host and guest separately; distinguish allocated RAM from resident use.
Minimum memory
Measure the viable guest/application memory floor under real load.
NUMA
Distinguish remote-node access and per-node pressure from total free RAM.

Worked example

Ten 16 GiB VMs on a 128 GiB host create 160 GiB nominal allocation, not proof of failure. Simultaneous 130 GiB resident demand may exceed capacity after host overhead. Correlate p95 latency with swap/PSI instead of uniformly shrinking all guests.

Example commands: replace lab values and confirm permissions and software versions before use.

virsh dommemstat lab-vm
vmstat 1 10
cat /proc/pressure/memory
numastat

Troubleshooting

ObservationLikely cause / distinctionVerification
Host swap risesResident demand and host overhead exceed capacity.Correlate vmstat and PSI over time.
Balloon statistics unavailableDriver/polling unsupported or disabled.Check guest driver and dommemstat coverage.

Acceptance checks

  1. Align host/guest measurements.
  2. Verify the balloon driver.
  3. Monitor swap-in/out rates.
  4. Inspect NUMA-node pressure.
  5. Determine guest minima with load tests.
  6. Remeasure p95 after changes.

Related concepts

Resource contention

Contention occurs when workloads sharing CPU, memory, storage or a link make each other wait. Apparent spare total capacity can hide a hot core or single-queue bottleneck. Correlate backup jobs, antivirus scans, index maintenance and user traffic on a common timeline. Confirm the bottleneck before adding resources. Run a workload alone and with its usual competitors to separate shared-resource effects, and report peak-hour latency alongside average utilization.

Capacity and usable headroom

Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.

Latency distribution

Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.

NUMA locality

On NUMA systems a processor may access nearby and remote-node memory at different costs. VM vCPU and memory sizes must be assessed against physical node capacity and hypervisor placement. More vCPUs do not always improve performance: reduced locality and scheduling waits can make it worse. Record physical topology, topology exposed to the guest and the workload’s memory behaviour together. Compare changes under the same load, using application latency and remote-memory effects rather than CPU percentage alone.

Telemetry and time correlation

Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.

Isolation versus virtualization

Containers and VMs provide different isolation boundaries. Containers share the host kernel; VMs run guest operating systems. Rootless execution, namespaces and capability restrictions can reduce risk, but configuration and host security still matter. Images, running containers and persistent volumes have separate lifecycles. Updating an image does not back up data. Verify process privileges, mounts, network access and persistence after recreation separately.

Primary documentation

Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.

Knowledge Center

Enterprise IT Product Sales, Licensing and Deployment
Enterprise IT Project & Solution Scenarios
View all related content
Text on WhatsApp
Copied!