Incident details

This document outlines the details of the incidents, related metrics, and the filters that you can apply.

Note: In the Controlled Availability release stage, Splunk products may have limitations on customer access, features, maturity, and regional availability. For additional information on Controlled Availability please contact your Splunk representative.

This Incidents dashboard provides a centralized view of incidents within Splunk Observability Cloud, enabling you to monitor, filter, and analyze system issues efficiently. It groups related alerts into single events to improve issue visibility.

Incident dashboard

Use the following filters to refine the list of incidents based on specific criteria:

Filter Description
State Filters incidents by their current lifecycle status, such as Active and Closed.
Severity Filters incidents by their assigned urgency level such as Critical and Major.
Entity Filters incidents associated with specific infrastructure components or services.
Environment Filters incidents based on the deployment environment.
Reset Clears all currently applied filters to show the default view.

Summary metrics

The dashboard displays four key performance indicators to provide a high-level overview of system health:

Metric Description
Active Incidents Displays the total number of incidents currently requiring attention.
Critical Incidents Displays the total number of incidents classified with a critical severity level.
Closed Incidents Indicates the total number of resolved incidents.
Average Resolution Calculates the mean time taken to resolve incidents.

List of incidents

The table lists individual incidents, providing granular details for each event. Use the search bar above the incident detail table to locate specific incidents by name.

Column Description
Status Indicates whether the incident is currently Active or Closed.
Severity Displays the urgency level, such as Critical, Major associated with the incident.
Incident Name Lists the descriptive name of the incident, often linked to the specific service or component affected.
Affected Entities Identifies the specific resources, nodes, or services impacted by the incident.
Alert Count Displays the number of individual alerts grouped into this incident.
Created Displays the timestamp indicating when the incident was first triggered.
Modified Displays the timestamp of the most recent update to the incident.
Duration Displays the total time elapsed since the incident was created or its current state.
Settings Provides access to additional configuration or management options for the specific row.

Incident summary

Select an incident from the list to view its details.

Toggle between the Overview tab and the Alerts tab to view the summary of the incident or the list of specific underlying alerts.

Incident summary

The Overview tab contains the detail information for the incident:

Section Description
Incident Summary Provides a detailed description of the event, including the trigger timestamp and the specific resource affected.
Impact Summary Identifies the scope of the issue, specifically listing impacted services or infrastructure entities, such as K8s Nodes or APM Services.
Alerts Lists the individual alerts that have been correlated and grouped into this specific incident. Use View All Alerts to see the full list.
Incident Timeline Tracks the chronological history of the incident, including creation time, status changes, and user or system actions. This serves as an audit trail for troubleshooting.

The Alerts tab provides a detailed breakdown of the specific alerts associated with an incident, allowing you to manage and investigate individual triggers.

The following metrics to help you quickly assess the severity and duration of the alerts:

Metric Description
Critical alerts Displays the total count of alerts currently classified as critical.
Alerts open >10m Indicates the number of alerts that have remained open for more than 10 minutes.

Use the tools in this section to search, filter, and review the list of alerts linked to the incident.

  • Add alerts: Select the button to manually associate additional alerts with the current incident.

  • Search: Use the search bar to find specific alerts by name.

The alert list provides the following details:

Column Description
Alert name Displays the descriptive name of the triggered alert.
Entity Identifies the specific host, node, or service associated with the alert.
Timestamp Displays the date and time the alert was triggered.
Duration Indicates how long the alert has been active.
Options Click the menu icon (three-dot) to access additional management actions such as open the alert or remove the alert from the incident.

Incident notification rules

You can create incident notification rules to define who receives notifications for incidents affecting specific services, infrastructure, or environments. For more information, see Incident notification rules.