Troubleshooting · Storage and Backup
Diagnose ZFS Checksum Errors: Disk, Cabling and Data
A checksum mismatch means a read block differs from its recorded verification value. Investigate disks, cables, controllers, memory and power. ZFS can repair from a healthy redundant…
Technical review:
Architecture and operating model
A checksum mismatch means a read block differs from its recorded verification value. Investigate disks, cables, controllers, memory and power. ZFS can repair from a healthy redundant copy; checksums on a single disk cannot recreate lost data.
Preserve zpool status -v, kernel logs and disk serial identities before clearing counters. Distinguish READ/WRITE/CKSUM from permanent errors. Scrub checks allocated data but cannot recover every damaged file without another valid copy.
Simultaneous errors across disks suggest a shared backplane/HBA/power path. Verify independent backups before force-imports, feature changes or device removal. Map TrueNAS pool identities to physical disk labels.
- 1Checksum finding
- 2Disk/shared-path logs
- 3Healthy copy or backup
- 4Scrub/file validation
Design parameters
- Evidence
- Preserve time, serial number, device path and event log together.
- Redundancy
- Check healthy copies at vdev level; spare capacity elsewhere does not protect that vdev.
- Repair verification
- After fixing hardware, use scrub and real file reads to check that errors do not recur.
Worked example
If two disks show CKSUM increases and SATA link resets in the same minute, inspect shared cable/power paths before replacing disks randomly. Restore damaged files from a known-good backup and compare application hashes. Monitor both a clean scrub and nonrecurrence of hardware errors.
Example commands: replace lab values and confirm permissions and software versions before use.
zpool status -v
zpool events -v
journalctl -k --since "1 hour ago"
# Schedule scrub after evidence capture and an operational impact assessment.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Errors on one disk | Disk or connection path. | Correlate SMART, kernel logs and serial number. |
| Permanent errors persist | Unrepairable block or no healthy copy. | List damaged files and scrub again after backup restoration. |
Acceptance checks
- Preserve evidence before clear.
- Verify device serial numbers.
- Inspect shared hardware paths.
- Read back a known-good backup.
- Record scrub results.
- Monitor recurring errors.
Related concepts
Durability and power loss
Survival of committed data depends on write guarantees across the database, OS, filesystem, controller and disks. Do not disable safety settings for speed without understanding caches and flush behaviour. Power-loss-protected storage helps but does not prove correctness of the entire chain. Run failure and recovery tests in a controlled lab, not production. Validate application records and supported consistency checks rather than merely opening a sample file.
Filesystem and storage layers
An application directory, mount point, logical volume and physical device belong to different layers. Inode exhaustion, read-only mounts and filesystem errors can stop writes even when space appears available. A snapshot often shares the same storage failure domain and is not an independent backup. Mount options and expansion methods depend on the filesystem. Verify mounts after reboot: writing to an unmounted directory can fill the wrong device.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
Memory errors and ECC
ECC helps detect and correct certain memory errors; it does not address every error type or hardware failure. Support depends on CPU, motherboard and firmware as well as the module. Rising correctable-error counts can indicate a maintenance need; uncorrectable errors can interrupt service. Review hardware logs with DIMM-slot and time information. Verify platform-specific population and module-mixing rules. Treat capacity planning and reliability planning separately: ECC does not replace backup or HA.
Replication scope
Replication transfers data or state changes to another copy. Synchronous transfer can introduce latency and connectivity dependence; asynchronous transfer can lag behind. A current replica is not necessarily protected from incorrect changes: deletion and corruption can propagate too. Compare lag with the application’s last committed operation, not merely connection status. Define which copy receives write authority after the source is lost and how the old node is reconciled when it returns. Report replication lag and the actual recoverable point separately.
Availability versus recovery
High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.