TrueNAS Performance and Maintenance: IOPS, ARC, Scrubs and Disk Health
Measure bottlenecks across clients, networking and disks; manage caches, scrubs, resilvers and updates alongside data safety.
Content update:
Architecture and operating model
Define block size, read/write ratio, random/sequential access, sync writes, concurrency and latency targets. MB/s alone does not describe user experience or datastore acceptance. Inspect p95/p99 latency with queues and CPU; averages can hide spikes. Use sample data and schedule disruptive benchmarks in change windows.
ARC is RAM cache, L2ARC additional read cache, and SLOG a separate synchronous-write intent-log device. SLOG is not a universal write accelerator; L2ARC is not an automatic fix for insufficient RAM. Measure hit ratio, working set and CPU/disk/network limits before adding cache. Disabling sync safety to improve benchmarks can violate durability expectations.
Correlate disk latency, queues, checksum errors, temperatures and device health. Cable/backplane failures can mimic disk faults; record timestamps, serial numbers and events before clearing counters. Scrubs read allocated data and verify checksums, repairing when suitable healthy redundant copies exist. Clean health counters alone do not prove recoverability.
Resilvering rebuilds replaced/failed-device data according to topology; it differs from scrubbing. Large disks and busy workloads extend duration. Avoid overlapping scrubs, resilvers, backups and scans during peak demand. Confirm bays/serials before replacement; removing the wrong healthy disk can exceed redundancy. Check pool status and alert delivery afterwards.
Measure errors, drops, MTU mismatches, TCP retransmissions and client limits. LACP does not automatically give one TCP flow the sum of link speeds; multipath and SMB multichannel are separate mechanisms requiring client/version/network validation. Verify jumbo frames end to end. A 10 GbE label does not guarantee 10 Gb/s file transfers.
Operations cover daily alerts, capacity trends, failed protection tasks, disk health and periodic restore verification. Updates need configuration backup, release notes, driver/API changes, dependent clients and a maintenance window. Repeat the same acceptance workload after changes; boot success alone is not maintenance acceptance.
- 1Workload and baseline
- 2End-to-end measurement
- 3Controlled change
- 4Acceptance and monitoring
Design parameters
- Latency distribution
- Measure p95/p99 and workload pattern alongside averages.
- Cache role
- Evaluate ARC, L2ARC and SLOG against distinct requirements.
- Data health
- Correlate scrub, resilver, checksum and device events.
- Maintenance overlap
- Schedule heavy tasks against workload interruption tolerance.
Worked example
Assume illustrative p95 latency rises from 8 ms to 80 ms when backup and scrub overlap. Keeping clients/data constant, separate tasks first, then inspect disk queues, network drops and CPU. Consider L2ARC only when evidence shows a cache need; these numbers are illustrative.
A 10 Gb/s link has a theoretical 1.25 GB/s ceiling; protocol, CPU, disks and clients reduce real throughput. At 250 MB/s, an illustrative 500 GB restore needs about 2,000 seconds for transfer alone, plus startup, verification and acceptance for RTO. Large-file speed may not represent small-file recovery.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| Copies are fast but applications slow | Small-I/O or sync latency may differ. | Measure p95/p99 under representative application I/O. |
| Checksum errors increase | Disk, cable or backplane faults may exist. | Inspect events by serial number and path. |
| Added cache makes no difference | Networking, CPU or writes may dominate. | Validate the investment with hit ratios and end-to-end metrics. |
Acceptance checks
- Compare measurements using the same workload and acceptance criteria.
- Measure IOPS, throughput and latency together.
- Assess caches by their distinct roles.
- Manage scrub/resilver and backup overlap.
- Confirm replacement disks by serial number.
- Test restores and alert delivery after maintenance.
Related concepts
Latency distribution
Latency is the time between starting an operation and receiving its result. An average can hide a small number of very slow operations; medians and p95/p99 percentiles answer different questions. Network RTT, storage waits, processor queues and application processing contribute to end-to-end time. State whether measurements come from the client or server. Check whether increases coincide with traffic growth, maintenance or capacity limits. Record normal and peak-hour baselines before selecting an alert threshold.
IOPS and block size
IOPS measures operations per second, whereas MB/s measures transferred data. Identical IOPS values produce very different bandwidth at different block sizes. For example, 10,000 operations/s at 4 KiB is approximately 39.1 MiB/s. Read/write mix, random versus sequential access and protection mechanisms affect the result. Measure actual block distributions and concurrency rather than transferring a benchmark result directly to production. High disk IOPS does not prove good application response; latency must be examined alongside it.
Bandwidth and useful throughput
Link capacity differs from useful application throughput. Protocol headers, encryption, retransmissions, small files and storage waits reduce net transfer speed. Keep bits and bytes distinct: 1 Gbit/s corresponds to a theoretical 125 MB/s, not an application performance guarantee. Estimate transfer time as data size divided by measured useful throughput. Observe the network, source reads and destination writes together to locate the bottleneck. Consider temporary slowdowns and competing workloads as well as average speed.
Queues and concurrency
A queue holds work arriving faster than a resource can process it. More concurrency can improve utilization up to a point, then increase waiting time. In a stable system Little’s law relates L = λ × W: average work in the system equals throughput multiplied by average total time. Units must agree. At 2,000 operations/s and 5 ms total time, roughly 10 operations are present concurrently. This is a planning relationship; queue limits, bursts and highly variable service times still require measurement.
Durability and power loss
Survival of committed data depends on write guarantees across the database, OS, filesystem, controller and disks. Do not disable safety settings for speed without understanding caches and flush behaviour. Power-loss-protected storage helps but does not prove correctness of the entire chain. Run failure and recovery tests in a controlled lab, not production. Validate application records and supported consistency checks rather than merely opening a sample file.
Write endurance and wear
SSD selection requires more than capacity and initial performance. Consider TBW, DWPD, warranty duration, power-loss protection and workload write patterns together. DWPD calculations must use the capacity and conditions defined by the manufacturer. Database logs, small random writes and temporary work files affect wear differently. Monitor SMART or vendor health data, considering temperature, media errors and remaining endurance together. Write amplification means application-written bytes may differ from the amount written to NAND.
Primary documentation
Editorial method
Prepared by the Doz Technology technical team using official vendor documentation. Scenarios and calculations illustrate the method; they are not completed customer tests. Before implementation, verify versions, licences, client support, security conditions and rollback in your environment. Select solutions against existing products and workload requirements.
Related storage guides
TrueNAS Installation: Hardware, Versions and Safe Migration
Plan the storage role, hardware compatibility, management access and version migration alongside recovery.
Read the guideTrueNAS and ZFS: Pools, Vdevs, RAIDZ and Dataset Design
Separate usable capacity from raw disks; plan mirror/RAIDZ topology, dataset boundaries, snapshot growth and expansion together.
Read the guideTrueNAS File and Storage Servers: SMB, NFS, iSCSI and ACLs
Separate file and block access; design access through identities, groups, share policies and client validation.
Read the guideTrueNAS: Snapshots, Replication and Backup Recovery
Design ZFS snapshots, local/remote replication and independent backups as distinct protection layers; measure RPO/RTO through restore tests.
Read the guideTrueNAS, OpenMediaVault, XigmaNAS and Unraid Comparison
Compare NAS choices by data model, disk topology, hardware compatibility, backup, operations and total cost.
Read the guideWe assess platform choice alongside your existing products, workloads and operational needs.