TrueNAS: Snapshots, Replication and Backup Recovery
Design ZFS snapshots, local/remote replication and independent backups as distinct protection layers; measure RPO/RTO through restore tests.
Content update:
Architecture and operating model
A snapshot preserves a dataset/zvol point in time without initially recopying all data. Retained changed blocks consume space. Same-pool snapshots are not independent copies against pool loss, nor automatically immutable against compromised storage administration. Useful local file recovery does not make a complete disaster-recovery plan.
Replication transfers snapshots to another dataset, pool or system. A different pool in the same chassis may separate some disk risk but still share fire, power, administration and location risk. Plan remote identity/access/retention boundaries. Record which side initiates push/pull with which privileges; remote location alone does not prevent attack propagation.
A backup-software NAS repository needs protocol/support validation, locking, write/read performance, independent copies and protection policies. An accessible SMB/NFS target is not automatically a Linux hardened repository or WORM store. Immutability must be implemented by the selected product/version/retention lock and verified through deletion tests.
Filesystem snapshots alone do not prove application consistency. Databases, VMs and mail data may need quiescing, application backup APIs or supported coordination. Transaction logs and restore order follow application requirements. Booting from a crash-consistent copy does not prove every transaction or business rule is intact.
Check retention, snapshot filters, recursive coverage and destination deletion together. Initial full and incremental transfers have different bandwidth needs. Measure the latest usable destination snapshot, not merely job success time. A successful task may still transfer stale or incomplete scope. Synchronise clocks and test connectivity/capacity alert delivery.
Restore to an isolated target where possible; production rollback can discard newer changes. Verify content, ACLs, ownership, application startup, dependencies and user acceptance together. Authorised administrators test encrypted-copy key availability. Manage configuration/key loss separately from data-copy availability.
- 1Scope and RPO/RTO
- 2Separate copies and privileges
- 3Isolated restore
- 4Acceptance and report
Design parameters
- Protection boundary
- State shared risks across snapshots, chassis and remote targets.
- Latest usable point
- Measure RPO against verified destination data time.
- Application consistency
- Validate application restore alongside filesystem state.
- Retention and deletion
- Test source/destination deletion and independent protection.
Worked example
If the latest usable remote snapshot is 10:40 and the incident is 11:00, the recovery-point gap is 20 minutes. Acceptance at 11:50 gives a 50-minute RTO. A 10:59 job transferring the 10:40 snapshot does not give a one-minute RPO. These are illustrative calculations, not service-time guarantees.
In a pilot, delete a file, change an ACL and update sample application data. Accept the required version, permissions and application transaction on an isolated restore target. Separately test whether source administration can delete the protected copy and whether lost destination access triggers an alert.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Replication succeeds but data is stale | Snapshot filters or source schedules may be wrong. | Check destination timestamps and recursive scope. |
| Pool is healthy but restore is incomplete | Child datasets, logs or dependencies may be omitted. | Match coverage against application acceptance criteria. |
| A protected copy can be deleted | Protection may only be snapshots or read-only access. | Test actual retention enforcement and administrative access. |
Acceptance checks
- Validate copies through actual recovery and access boundaries.
- Distinguish snapshots from independent backup.
- Test destination retention and deletion.
- Verify application consistency and scope.
- Record the latest usable data point.
- Complete restore with ACL and user acceptance.
Related concepts
Replication scope
Replication transfers data or state changes to another copy. Synchronous transfer can introduce latency and connectivity dependence; asynchronous transfer can lag behind. A current replica is not necessarily protected from incorrect changes: deletion and corruption can propagate too. Compare lag with the application’s last committed operation, not merely connection status. Define which copy receives write authority after the source is lost and how the old node is reconciled when it returns. Report replication lag and the actual recoverable point separately.
RPO and the actual loss window
RPO is the acceptable duration of data loss after an incident. A backup schedule alone does not prove it: failed jobs or delayed replication to another site can enlarge the window. Compare the incident time with the latest usable, consistent recovery point. If an incident occurs at 14:00 and the verified copy is from 13:20, the observed loss window is 40 minutes. Set targets per application; a file archive and a database receiving continuous orders may have different requirements.
RTO and end-to-end recovery time
RTO specifies how soon a service must become usable after an incident. Downloading a backup or booting a VM accounts for only part of that duration. Record detection, approval, infrastructure preparation, data transfer, application startup and business validation separately. Define exactly what starts and stops the test clock. The same technical restore duration can produce different business interruptions when dependencies or access approvals introduce additional waiting.
Immutability and deletion authority
Immutable retention restricts modification or deletion for a defined period. An application promise, storage-enforced protection and physical isolation are different controls. Review lock duration, clock handling, administrative privileges and deletion paths together. Whether data is automatically deleted after expiry depends on the product policy. On a controlled test copy, attempt deletion using ordinary and privileged accounts, then verify recovery from the protected copy. Protection against deletion does not replace a readability or restore test.
Retention and capacity
Retention defines which recovery points are kept and for how long. Daily, weekly and monthly points do not represent identical change patterns; full-copy creation and chain dependencies affect physical capacity. Retention decisions combine business requirements, applicable obligations and technical capacity. Longer retention does not automatically provide better recovery: the right point must be discoverable and readable. When changing a policy, test whether existing points are deleted immediately or handled differently by the product.
Application consistency
A copy that boots does not prove application-data consistency. Operating-system caches, database logs and write ordering across disks or services affect the result. A crash-consistent copy resembles recovery after an unexpected shutdown; an application-consistent copy follows supported application preparation and write coordination. Validate transaction integrity, relationships between records and application behaviour after recovery, rather than only counting files. Confirm backup integration, application version and reported errors before assuming that consistency was achieved.
Primary documentation
Editorial method
Prepared by the Doz Technology technical team using official vendor documentation. Scenarios and calculations illustrate the method; they are not completed customer tests. Before implementation, verify versions, licences, client support, security conditions and rollback in your environment. Select solutions against existing products and workload requirements.
Related storage guides
TrueNAS Installation: Hardware, Versions and Safe Migration
Plan the storage role, hardware compatibility, management access and version migration alongside recovery.
Read the guideTrueNAS and ZFS: Pools, Vdevs, RAIDZ and Dataset Design
Separate usable capacity from raw disks; plan mirror/RAIDZ topology, dataset boundaries, snapshot growth and expansion together.
Read the guideTrueNAS File and Storage Servers: SMB, NFS, iSCSI and ACLs
Separate file and block access; design access through identities, groups, share policies and client validation.
Read the guideTrueNAS Performance and Maintenance: IOPS, ARC, Scrubs and Disk Health
Measure bottlenecks across clients, networking and disks; manage caches, scrubs, resilvers and updates alongside data safety.
Read the guideTrueNAS, OpenMediaVault, XigmaNAS and Unraid Comparison
Compare NAS choices by data model, disk topology, hardware compatibility, backup, operations and total cost.
Read the guideWe assess platform choice alongside your existing products, workloads and operational needs.