Signs of a drive failure
Drive failures typically surface through a combination of application-level and storage-level symptoms:- Write errors in the admin UI: Services report that they cannot write data, or storage-related warnings and errors appear in the system status view.
- Services in failed or crash-looping states: Applications that require storage access may repeatedly fail to start or exit immediately after starting.
- Missing or incomplete data in analyst tools: Packet captures stop updating, telemetry shows gaps, or the flow analytics view becomes stale.
- Explicit storage fault messages: The BMC or OS-level logs visible through the BMC console may report SMART errors, RAID degradation (on models with redundant storage), or I/O errors.
Immediate steps
1
Stop accepting new captures if possible
If the appliance is still partially operational, consider pausing new traffic ingestion to reduce the rate of new write failures. This limits further data integrity issues while you assess the situation.
2
Attempt to offload accessible data
If any data remains accessible through the analyst tools or the data management interface, initiate a data offload to export it to an external destination before the failure progresses. See Data Offloading. Not all data may be recoverable depending on the extent of the failure.
3
Document what you are seeing
Note the specific error messages visible in the admin UI, the system status view, and (if accessible) the BMC console. Capture screenshots if possible. This information will be necessary when you contact support.
4
Escalate to support
Contact support with the information you have gathered. See When to Escalate for what to include. Drive replacement and data recovery at the hardware level require vendor intervention.
On APEX units with multiple nodes, a failure on one node may not immediately affect the other nodes. Check which node is reporting errors, as the remaining nodes may continue to operate normally while the affected node is serviced.