Configure the Prometheus receiver to collect NetApp metrics
Configure the Prometheus receiver to collect NetApp metrics.
You can monitor the performance of NetApp storage resources the following NetApp tools:
-
NetApp Trident: A storage orchestrator and management tool for containers and Kubernetes distributions.
-
NetApp Harvest: An open-metrics endpoint for ONTAP, StorageGRID, E-Series and Cisco Nexus switches.
Splunk Observability Cloud uses the Prometheus receiver to collect metrics from NetApp Trident and Harvest, which expose /metrics endpoints that publish Prometheus-compatible metrics.
This data integration requires:
-
A Kubernetes or OpenShift cluster with NetApp Trident installed.
-
Trident metrics exposed from the trident-csi service. Trident exposes Prometheus metrics at /metrics, typically on port 8001.
-
NetApp Harvest deployed and configured to collect ONTAP metrics from NetApp clusters.
-
NetApp Harvest configured with ONTAP cluster connection details and appropriate ONTAP API credentials.
Configuration settings
Learn about the configuration options for the Prometheus receiver.
To view the configuration options for the Prometheus receiver, see Settings.
Metrics
NetApp storage monitoring in Splunk Observability Cloud uses two main metric sources:
-
NetApp Harvest: Used for ONTAP cluster, controller, aggregate, volume, latency, throughput, network, and hardware/environment telemetry.
-
NetApp Trident: Used for Kubernetes/OpenShift CSI orchestration telemetry such as volume count, operation count, latency, failures, and backend health.
The following metrics are available for NetApp. For more information about these metrics, see Monitor Trident and ONTAP Metrics in the NetApp documentation.
These metrics are considered custom metrics in Splunk Observability Cloud.
| Metric name | Description |
|---|---|
trident_backend_count |
The total number of backends. |
trident_node_count |
The total number of nodes. |
trident_operation_duration_milliseconds_count |
The total count of observed operations. |
trident_operation_duration_milliseconds_quantile |
The latency quantile for operation events. |
trident_operation_duration_milliseconds_sum |
The total duration of all observed operations. |
trident_storageclass_count |
The total number of storage classes. |
cluster_new_status |
The current cluster status. |
cluster_subsystem_outstanding_alerts |
The current unresolved subsystem alert count. |
node_avg_processor_busy |
The average processor busy percentage. |
Next steps
After you set up data collection, the data populates built-in dashboards that you can use to monitor and troubleshoot your instances.
For more information on using built-in dashboards in Splunk Observability Cloud, see: