Monitoring and alerts
Object Storage provides out-of-the-box monitoring metrics collection and alerting capabilities, allowing you to monitor and alert on storage clusters, storage services, and configure notification policies.
Intuitive monitoring data can be used to provide decision support for operations and performance tuning, and a comprehensive alert and notification mechanism will help ensure the stable operation of the storage system.
Monitoring
The platform collects the health status of the storage cluster and storage services by default. On the Monitoring tab of Storage > Object Storage, you can view real-time monitoring data for metrics.
Storage overview
Monitor the health status of storage, storage service status, and cluster raw capacity usage. When the storage status is abnormal, you can view the cause of the alert.
Cluster monitoring
Monitor the raw capacity usage and read/write rate of the storage cluster.
Object monitoring
Monitor the total number of access requests and error access requests for objects.
Alerts
The platform enables a set of alert policies by default. Once a resource is abnormal or monitoring data reaches a warning state, an alert will be triggered automatically. The preset policies can meet common operational requirements, such as component and cluster status alerts, device capacity alerts, and user data alerts.
Configuration notification
To receive alerts in a timely manner, it is recommended that you set up notification policies in the Operations Center. This will send alert information to relevant personnel via email, SMS, etc., reminding them to take necessary measures to solve problems or avoid faults. Click on to switch to the Operations Center to complete the operation, and refer to
Creating Alert Policies
for guidance.
Handle alerts
-
If the storage cluster is in an “Alarm” state, it means that an alert has been triggered and related anomalies may cause faults. Please check the details of Realtime Alerts in a timely manner and locate and troubleshoot the problem based on the cause of the fault.
-
If the storage cluster is in a “Fault” state, it means that the storage cluster is no longer functioning properly. Please locate and troubleshoot the problem immediately.
The following table shows the meaning of the alarm levels used in the preset policies, which can serve as a reference for you to formulate alarm handling principles.
| Alarm Level | Meaning |
|---|---|
| Critical | The resource corresponding to the alarm rule has failed, causing a significant impact on platform business interruption and data loss. |
| Height | The resource corresponding to the alarm rule has a known issue that may cause platform function failure, affecting normal business operations. |
| Medium | The resource corresponding to the alarm rule has operational risks. If not handled in a timely manner, it may affect normal business operations. |
Fault review
History records all alerts that have been triggered and no longer require attention. When using alert history for fault replay, in order to effectively achieve the goal of experience summary, you may need to answer the following questions.
- What were the specific abnormal situations when the accident occurred?
- Is there any pattern to the repeated appearance of a certain alarm in the alarm list? Can it be avoided before the next occurrence?
- If the time axis shows a surge in alarms during a certain period, is it due to force majeure or operational accidents? Is it necessary to adjust the operational plan?