Degraded operation

Estimate a NAS rebuild window before a drive fails

A rebuild plan is most useful before an incident, when you can measure the storage path, choose a workload policy, and verify the independent recovery route.

Start with the reconstruction scope

Use the amount of data the platform expects to reconstruct or verify. When that value is not yet available, the replacement drive label provides a visible upper planning boundary that can be refined after a scrub, repair, or resilver measurement.

The planner uses decimal units consistently: one entered TB becomes 1,000,000 MB for the time calculation. Keeping that conversion visible prevents a unit mismatch from being mistaken for storage overhead.

Model the path, not a drive-speed headline

A rebuild must read from healthy members, pass through the controller or storage pipeline, and write to the replacement member. The slowest entered rate sets the raw transfer bottleneck. This is more useful than assigning one speed to a drive model because the limiting stage can be elsewhere in the system.

Reserve a share of that bottleneck for foreground services, then apply a planning-efficiency factor for scheduling, seeks, retries, and implementation behavior. The result is an editable planning rate rather than a hidden multiplier.

Keep the range visible

The transfer floor divides the reconstruction scope by the raw bottleneck. The planning window divides the same scope by the load- and efficiency-adjusted rate. Compare the planning edge with the maximum degraded duration your storage and recovery plan can accept.

If the window is too long, the useful decisions are concrete: measure the slowest stage, reduce competing workload during repair, reduce the critical data dependency, or improve the independent recovery path.

Read fault tolerance in the degraded state

Single-parity and two-drive mirror paths have no additional drive failure inside their simplified tolerance after one member fails. Dual-parity paths retain one additional drive-failure boundary. RAID 10 depends on which mirror pair contains a subsequent failure, so the exact member layout remains part of the incident procedure.

Rehearse the response boundary

Record how the platform identifies the failed member, which replacement sizes and types it accepts, how repair starts, where progress and errors appear, and what workload policy applies until the array returns to a healthy state. Keep that procedure beside a measured restore path, because degraded-array availability and independent recovery solve different problems.

Continue planning

Turn this storage question into a capacity baseline.

Use the planner to make growth, bays, redundancy, and recovery capacity visible together.

Editorial record

Maintained by Make Your Own Tools to turn “Store data safely” into a defined capacity, bays, redundancy, and recovery baseline. The references below provide the technical context for this planning path. Its calculation rules and planning assumptions are documented in the methodology, and affected calculations pass regression checks before the review date advances.

Last reviewed
Evidence set
2 primary references
Calculation coverage
1 affected rule checked
Planning scope
Pre-incident planning for a protected NAS array that can repair, rebuild, or resilver onto a replacement drive.
Next check
Replace the planning inputs with observed platform rates and verify drive requirements, alerts, workload policy, and independent recovery procedures.