Overview
The Operations Center provides monitoring, logging, event, alarm, and notification capabilities for platform administrators and operations administrators.
The platform integrates with Prometheus monitoring and Grafana visualization panels, supporting real-time monitoring of clusters, nodes, components, custom applications, Pods, and containers managed by the platform.
It supports quick setting of monitoring metric alarms for clusters, nodes, and computing components, log alarms (for computing components only), and event alarms (for computing components only). Users can also customize monitoring metric algorithms and add required alarm metrics and rules based on actual needs. By configuring notification policies, alarm information can be sent to operations personnel in a timely manner to avoid system failures or promptly handle faults, reduce system operation and maintenance costs, and ensure system stability.
At the same time, the platform fully integrates Kubernetes events, allowing users to view all Kubernetes event information on the platform. It also supports viewing the log information of all resources on the platform, including the standard output logs of containers and the logs recorded in specified files within containers, which can help users quickly troubleshoot and resolve issues.
Switching of clusters
The monitoring panel, alarms, and events are cluster-level resources. Before using these functions, you need to switch to a specific cluster, as shown in the figure below.