Skip to main content

Comparison · Linux and System Administration

Cgroup v2 MemoryHigh versus MemoryMax

memory.high applies pressure/reclaim and can throttle a workload above the threshold; it is not a hard usage ceiling. memory.max is a hard limit and can lead to cgroup OOM when reclaim…

Technical review:

Architecture and operating model

memory.high applies pressure/reclaim and can throttle a workload above the threshold; it is not a hard usage ceiling. memory.max is a hard limit and can lead to cgroup OOM when reclaim fails. They are not interchangeable variants.

Parent limits may constrain nested cgroups. An 8 GiB container/service limit cannot override a 6 GiB parent constraint. Record swap settings and kernel version; RSS does not represent all charged memory.

Setting MemoryHigh too low without a baseline can worsen latency. Correlate memory.events high/max/oom/oom_kill with p95 and PSI. Ensure restart policy does not create an OOM loop.

  1. 1Memory allocation
  2. 2High pressure/reclaim
  3. 3Max hard limit
  4. 4Counters/application impact
Read parent limits.

Design parameters

High threshold
Set above the normal working set through a pilot; constant reclaim is not the goal.
Max limit
Protect the host while defining accepted workload overflow/failure behavior.
Counters
Account for counter lifecycle across service restarts.

Platform implementation

Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.

ControlBehaviorObservation
memory.highPressure/reclaim; not a hard caphigh + latency/PSI
memory.maxHard limit, reclaim/OOMmax/oom/oom_kill
Parent limitConstrains descendantsEntire hierarchy

Worked example

For a lab service using 2 GiB normally and 3 GiB briefly, High=2500M/Max=4G can be piloted, not universally recommended. Rising high counts and p95 indicate pressure. Measure failure/restart behavior under a controlled max-limit test.

Troubleshooting

ObservationLikely cause / distinctionVerification
Slow service below maxHigh reclaim or parent pressure.Compare memory.events, parent limits and PSI.
Repeated OOMMax too low or memory leak.Inspect working-set trend and restart events.

Acceptance checks

  1. Read parent limits.
  2. Measure normal working set.
  3. Test high effects on p95.
  4. Test failure behavior at max.
  5. Verify swap policy.
  6. Prevent OOM restart loops.

Related concepts

Isolation versus virtualization

Containers and VMs provide different isolation boundaries. Containers share the host kernel; VMs run guest operating systems. Rootless execution, namespaces and capability restrictions can reduce risk, but configuration and host security still matter. Images, running containers and persistent volumes have separate lifecycles. Updating an image does not back up data. Verify process privileges, mounts, network access and persistence after recreation separately.

Resource contention

Contention occurs when workloads sharing CPU, memory, storage or a link make each other wait. Apparent spare total capacity can hide a hot core or single-queue bottleneck. Correlate backup jobs, antivirus scans, index maintenance and user traffic on a common timeline. Confirm the bottleneck before adding resources. Run a workload alone and with its usual competitors to separate shared-resource effects, and report peak-hour latency alongside average utilization.

Capacity and usable headroom

Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.

Service lifecycle and dependencies

A running process does not prove that users receive correct service. Review startup order, network or database dependencies, environment variables, file access and health checks together. Restart loops can conceal the root cause. Compare exit codes, recent logs and resource limits, and identify whether systemd or container management triggered the restart. After a controlled restart, validate sessions, writes and dependent services.

Latency distribution

Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.

Telemetry and time correlation

Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.

Primary documentation

Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.

Knowledge Center

Enterprise IT Product Sales, Licensing and Deployment
Enterprise IT Project & Solution Scenarios
View all related content
Text on WhatsApp
Copied!