Operations center
The Operations Center provides monitoring, logging, event, alarm, and notification capabilities for platform administrators and operations administrators.
The platform supports black box monitoring, allowing users to detect network issues quickly through HTTP, HTTPS, TCP, DNS, and ICMP probes. Additionally, the platform collects monitoring metrics and presents real-time monitoring data based on these metrics in a visual dashboard. Users can view the monitoring dashboard to understand or predict the health status of the platform's functionality.
The platform collects and saves system logs, product logs, Kubernetes logs, and custom application logs. Platform administrators or operations personnel can manage log retention policies, view categorized logs, and export logs through the logging module.
The platform integrates with Kubernetes events, recording important status changes of Kubernetes resources and various runtime state changes, and provides storage, querying, and visualization capabilities. When there are abnormal situations with clusters, nodes, Pods, and other resources, the specific reasons can be analyzed through event analysis.
Platform administrators can manage the alert policies for clusters, nodes, and computing components through the Operations Center. At the same time, they can centrally manage the alert templates for the platform, making it easy to create alert policies for clusters, nodes, and computing components.
Notifications support sending platform operation status, such as platform monitoring and alarm information, in the form of email, SMS, webhook, DingTalk, and WeChat Work.
Inspection can help enterprise customers understand the running status of business resources on the platform in real time, perceive anomalies in a timely manner, and reduce business risks.