Degraded operation
Estimate a NAS rebuild window before a drive fails
A rebuild plan is most useful before an incident, when you can measure the storage path, choose a workload policy, and verify the independent recovery route.
Start with the reconstruction scope
Use the amount of data the platform expects to reconstruct or verify. When that value is not yet available, the replacement drive label provides a visible upper planning boundary that can be refined after a scrub, repair, or resilver measurement.
The planner uses decimal units consistently: one entered TB becomes 1,000,000 MB for the time calculation. Keeping that conversion visible prevents a unit mismatch from being mistaken for storage overhead.
Model the path, not a drive-speed headline
A rebuild must read from healthy members, pass through the controller or storage pipeline, and write to the replacement member. The slowest entered rate sets the raw transfer bottleneck. This is more useful than assigning one speed to a drive model because the limiting stage can be elsewhere in the system.
Reserve a share of that bottleneck for foreground services, then apply a planning-efficiency factor for scheduling, seeks, retries, and implementation behavior. The result is an editable planning rate rather than a hidden multiplier.
Keep the range visible
The transfer floor divides the reconstruction scope by the raw bottleneck. The planning window divides the same scope by the load- and efficiency-adjusted rate. Compare the planning edge with the maximum degraded duration your storage and recovery plan can accept.
If the window is too long, the useful decisions are concrete: measure the slowest stage, reduce competing workload during repair, reduce the critical data dependency, or improve the independent recovery path.
Read fault tolerance in the degraded state
Single-parity and two-drive mirror paths have no additional drive failure inside their simplified tolerance after one member fails. Dual-parity paths retain one additional drive-failure boundary. RAID 10 depends on which mirror pair contains a subsequent failure, so the exact member layout remains part of the incident procedure.
Rehearse the response boundary
Record how the platform identifies the failed member, which replacement sizes and types it accepts, how repair starts, where progress and errors appear, and what workload policy applies until the array returns to a healthy state. Keep that procedure beside a measured restore path, because degraded-array availability and independent recovery solve different problems.