Comparison · Linux and System Administration
Cgroup v2 MemoryHigh versus MemoryMax
memory.high applies pressure/reclaim and can throttle a workload above the threshold; it is not a hard usage ceiling. memory.max is a hard limit and can lead to cgroup OOM when reclaim…
Technical review:
Architecture and operating model
memory.high applies pressure/reclaim and can throttle a workload above the threshold; it is not a hard usage ceiling. memory.max is a hard limit and can lead to cgroup OOM when reclaim fails. They are not interchangeable variants.
Parent limits may constrain nested cgroups. An 8 GiB container/service limit cannot override a 6 GiB parent constraint. Record swap settings and kernel version; RSS does not represent all charged memory.
Setting MemoryHigh too low without a baseline can worsen latency. Correlate memory.events high/max/oom/oom_kill with p95 and PSI. Ensure restart policy does not create an OOM loop.
- 1Memory allocation
- 2High pressure/reclaim
- 3Max hard limit
- 4Counters/application impact
Design parameters
- High threshold
- Set above the normal working set through a pilot; constant reclaim is not the goal.
- Max limit
- Protect the host while defining accepted workload overflow/failure behavior.
- Counters
- Account for counter lifecycle across service restarts.
Platform implementation
Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.
| Control | Behavior | Observation |
|---|---|---|
| memory.high | Pressure/reclaim; not a hard cap | high + latency/PSI |
| memory.max | Hard limit, reclaim/OOM | max/oom/oom_kill |
| Parent limit | Constrains descendants | Entire hierarchy |
Worked example
For a lab service using 2 GiB normally and 3 GiB briefly, High=2500M/Max=4G can be piloted, not universally recommended. Rising high counts and p95 indicate pressure. Measure failure/restart behavior under a controlled max-limit test.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Slow service below max | High reclaim or parent pressure. | Compare memory.events, parent limits and PSI. |
| Repeated OOM | Max too low or memory leak. | Inspect working-set trend and restart events. |
Acceptance checks
- Read parent limits.
- Measure normal working set.
- Test high effects on p95.
- Test failure behavior at max.
- Verify swap policy.
- Prevent OOM restart loops.
Related concepts
Isolation versus virtualization
Containers and VMs provide different isolation boundaries. Containers share the host kernel; VMs run guest operating systems. Rootless execution, namespaces and capability restrictions can reduce risk, but configuration and host security still matter. Images, running containers and persistent volumes have separate lifecycles. Updating an image does not back up data. Verify process privileges, mounts, network access and persistence after recreation separately.
Resource contention
Contention occurs when workloads sharing CPU, memory, storage or a link make each other wait. Apparent spare total capacity can hide a hot core or single-queue bottleneck. Correlate backup jobs, antivirus scans, index maintenance and user traffic on a common timeline. Confirm the bottleneck before adding resources. Run a workload alone and with its usual competitors to separate shared-resource effects, and report peak-hour latency alongside average utilization.
Capacity and usable headroom
Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.
Service lifecycle and dependencies
A running process does not prove that users receive correct service. Review startup order, network or database dependencies, environment variables, file access and health checks together. Restart loops can conceal the root cause. Compare exit codes, recent logs and resource limits, and identify whether systemd or container management triggered the restart. After a controlled restart, validate sessions, writes and dependent services.
Latency distribution
Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.