Technical Guide · Storage and Backup
Plan Bandwidth and Lag for ZFS Replication
ZFS incremental send transfers changes after a shared snapshot. Transfer size may differ from visible file changes because of blocks, metadata and compression modes. Data that has not…
Technical review:
Architecture and operating model
ZFS incremental send transfers changes after a shared snapshot. Transfer size may differ from visible file changes because of blocks, metadata and compression modes. Data that has not been applied and made accessible on the receiver should not count toward accepted RPO.
Size initial transfer, regular changes and post-outage backlog separately. If throughput remains below incoming change rate, lag never clears. Ensure TrueNAS task metrics and underlying ZFS transfer metrics refer to the same time window.
Compressed/raw encrypted send options must match receiver features and key management. Separate destination deletion authority from source identity; replication alone is neither independent nor immutable backup.
- 1Shared snapshot
- 2Changed-block stream
- 3Link and backlog
- 4Receiver apply/validation
Design parameters
- Effective throughput
- Use measured payload throughput rather than line rate, allowing for protocol and competing traffic.
- Backlog
- Calculate outage changes alongside the normal ongoing change rate.
- Snapshot dependency
- Removing the shared snapshot can break incremental continuity; coordinate retention policies.
Worked example
A daily change of 300 GB averages 3.47 MB/s. With 10 MB/s effective throughput, a 300 GB backlog clears theoretically in 300,000/(10−3.47) ≈ 45,942 seconds, about 12.8 hours. This decimal estimate excludes metadata and peak periods; add headroom from pilot measurements.
Example commands: replace lab values and confirm permissions and software versions before use.
zfs send -nP -i tank/data@base tank/data@next
zfs list -t snapshot -o name,used,creation
# Dry-run output estimates a particular stream; actual network throughput still needs measurement.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Lag continuously grows | Effective transfer slower than changes. | Compare bytes generated and transferred in the same window. |
| Incremental fails to start | Shared snapshot missing. | Inspect snapshot GUIDs and retention records on both sides. |
Acceptance checks
- Size the initial copy separately.
- Measure daily changes.
- Test backlog clearance after outage.
- Preserve the shared snapshot.
- Read data back from the receiver.
- Separate deletion permissions.
Related concepts
Replication scope
Replication transfers data or state changes to another copy. Synchronous transfer can introduce latency and connectivity dependence; asynchronous transfer can lag behind. A current replica is not necessarily protected from incorrect changes: deletion and corruption can propagate too. Compare lag with the application’s last committed operation, not merely connection status. Define which copy receives write authority after the source is lost and how the old node is reconciled when it returns. Report replication lag and the actual recoverable point separately.
Bandwidth and useful throughput
Link capacity differs from useful application throughput. Protocol headers, encryption, retransmissions, small files and storage waits reduce net transfer speed. Keep bits and bytes distinct: 1 Gbit/s corresponds to a theoretical 125 MB/s, not an application performance guarantee. Estimate transfer time as data size divided by measured useful throughput. Observe the network, source reads and destination writes together to locate the bottleneck. Consider temporary slowdowns and competing workloads as well as average speed.
RPO and the actual loss window
RPO is the acceptable duration of data loss after an incident. A backup schedule alone does not prove it: failed jobs or delayed replication to another site can enlarge the window. Compare the incident time with the latest usable, consistent recovery point. If an incident occurs at 14:00 and the verified copy is from 13:20, the observed loss window is 40 minutes. Set targets per application; a file archive and a database receiving continuous orders may have different requirements.
Capacity and usable headroom
Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.
Encryption and key lifecycle
Encryption makes data difficult to read without its key; access control, deletion protection and backup address different needs. Identify which layer protects data in transit and data at rest. Key generation, protection, rotation and emergency recovery are part of the design. A key needed to restore a backup must not exist only on the server that could be lost in the incident. Test decryption and key-access recovery using a separate administrator in a lab, and keep real keys out of documentation and support messages.
Retention and capacity
Retention defines which recovery points are kept and for how long. Daily, weekly and monthly points do not represent identical change patterns; full-copy creation and chain dependencies affect physical capacity. Retention decisions combine business requirements, applicable obligations and technical capacity. Longer retention does not automatically provide better recovery: the right point must be discoverable and readable. When changing a policy, test whether existing points are deleted immediately or handled differently by the product.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.