Comparison · Storage and Backup
ZFS Compression versus Dedup: Space Savings and Operational Cost
Compression represents data within a block more compactly; dedup shares identical-content blocks. Similar file names or many VMs do not prove a high dedup ratio. Precompressed media and…
Technical review:
Architecture and operating model
Compression represents data within a block more compactly; dedup shares identical-content blocks. Similar file names or many VMs do not prove a high dedup ratio. Precompressed media and encrypted data may also reduce compression benefits.
Dedup table access adds metadata work on writes and frees. New fast-dedup features have release/pool-feature requirements; do not assume identical behavior on older TrueNAS/OpenZFS installations. Changing compression/dedup does not automatically rewrite existing blocks.
- 1Representative blocks
- 2Compression
- 3Dedup table
- 4Space and I/O measurement
Design parameters
- Real data sample
- Use representative copied data rather than synthetic zero blocks.
- Metadata cost
- Measure RAM, metadata I/O and deletion time alongside savings.
- Transition
- Plan for pool features that may constrain rollback to older releases.
Platform implementation
Service names are reference points. Scope, defaults, region availability and operating requirements differ; they are not interchangeable guarantees.
| Feature | Compression | Dedup |
|---|---|---|
| Working unit | Within-block patterns | Identical-content blocks |
| Additional cost | CPU/algorithm | Table/metadata access |
| Pilot metric | Space + throughput | Incremental savings + metadata/recovery |
Worked example
Suppose 1 TB of representative data uses 600 GB with compression and 550 GB with compression plus dedup. Dedup saves an additional 50 GB, an 8.3% reduction relative to compression. Do not choose based solely on a ratio if metadata load and recovery time worsen.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| No space saving | Unique or precompressed data. | Inspect compressratio/dedupratio on representative samples. |
| Deletion is slow | Dedup metadata operations dominate. | Record disk latency and metadata/cache indicators. |
Acceptance checks
- Use representative data.
- Measure a compression-only baseline.
- Calculate incremental dedup savings.
- Monitor metadata I/O.
- Test deletion and restore times.
- Verify release/pool-feature compatibility.
Related concepts
Capacity and usable headroom
Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.
Resource contention
Contention occurs when workloads sharing CPU, memory, storage or a link make each other wait. Apparent spare total capacity can hide a hot core or single-queue bottleneck. Correlate backup jobs, antivirus scans, index maintenance and user traffic on a common timeline. Confirm the bottleneck before adding resources. Run a workload alone and with its usual competitors to separate shared-resource effects, and report peak-hour latency alongside average utilization.
Bandwidth and useful throughput
Link capacity differs from useful application throughput. Protocol headers, encryption, retransmissions, small files and storage waits reduce net transfer speed. Keep bits and bytes distinct: 1 Gbit/s corresponds to a theoretical 125 MB/s, not an application performance guarantee. Estimate transfer time as data size divided by measured useful throughput. Observe the network, source reads and destination writes together to locate the bottleneck. Consider temporary slowdowns and competing workloads as well as average speed.
IOPS and block size
IOPS measures operations per second, whereas MB/s measures transferred data. Identical IOPS values produce very different bandwidth at different block sizes. For example, 10,000 operations/s at 4 KiB is approximately 39.1 MiB/s. Read/write mix, random versus sequential access and protection mechanisms affect the result. Measure actual block distributions and concurrency rather than transferring a benchmark result directly to production. High disk IOPS does not prove good application response; latency must be examined alongside it.
Write endurance and wear
SSD selection requires more than capacity and initial performance. Consider TBW, DWPD, warranty duration, power-loss protection and workload write patterns together. DWPD calculations must use the capacity and conditions defined by the manufacturer. Database logs, small random writes and temporary work files affect wear differently. Monitor SMART or vendor health data, considering temperature, media errors and remaining endurance together. Write amplification means application-written bytes may differ from the amount written to NAND.
Filesystem and storage layers
An application directory, mount point, logical volume and physical device belong to different layers. Inode exhaustion, read-only mounts and filesystem errors can stop writes even when space appears available. A snapshot often shares the same storage failure domain and is not an independent backup. Mount options and expansion methods depend on the filesystem. Verify mounts after reboot: writing to an unmounted directory can fill the wrong device.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.