Health Dashboard

Use the On Premise Health Dashboard in the Virtual Appliance Platform UI to monitor recent cluster and service health, license usage, and hourly module health.

Use the On Premise Health Dashboard in the Virtual Appliance Platform UI to monitor recent cluster and service health, license usage, and hourly module health. The dashboard refreshes health, license, and heatmap data every 30 seconds.

Table 1. Health Statuses
Status Description
Healthy The component is functioning as expected.
Degraded The component is operational but has one or more warnings.
Unhealthy A critical health check has failed.
Unknown The component's health could not be determined because health information is unavailable.

Open the Health Dashboard

Open the Health Dashboard from the Virtual Appliance Platform UI.

Obtain credentials for the Virtual Appliance Platform UI.
  1. Log in to the Virtual Appliance Platform UI.
  2. Open /on-premise-health-dashboard on the same Virtual Appliance host.
  3. Wait for all four panels to load.
    The page displays Infrastructure Status, Splunk AppDynamics Services, Deployment Activity, and Health Heatmap.
The Health Dashboard displays the latest available health and license information and refreshes it every 30 seconds.

Health Dashboard panels

Use the four dashboard panels to investigate infrastructure, services, license usage, and historical module health.

Infrastructure Status

The Infrastructure Status panel displays overall cluster status, node count, and update time. Nodes with more unhealthy pods appear first. The remaining order uses health severity, from Unhealthy through Healthy, and then node name.

Expand a node to review:

  • CPU usage, memory usage, root file system usage, AppDynamics logs file system usage, and AppDynamics data file system usage.
  • CPU cores, total RAM, node uptime, and node IP address.
  • Node issues reported by the infrastructure health API.
  • Unhealthy pod names, failure reasons, and restart counts.
Gauge File system path
Root FS Used /
Logs FS Used /var/appd/logs
Data FS Used /var/appd/data

Interpret status according to the following thresholds:

Gauge Healthy Degraded Unhealthy
CPU 80% or less More than 80% through 95% More than 95%
Memory and file systems 85% or less More than 85% through 95% More than 95%

Use reported issues and pod failures together with the utilization values when investigating node health.

Infrastructure status and resource usage for cluster nodes.

Splunk AppDynamics Services

The Splunk AppDynamics Services panel provides a drill-down from modules to services, pods, and component dimensions.

When you open a service, the panel initially selects an Unhealthy pod, if present. Otherwise, it selects a Degraded pod or the first pod.

Pod details can include:

  • Version, namespace, last-polled time, and uptime when available.
  • Liveness and readiness probe status.
  • Overall pod status and the number of healthy components.
  • Symptoms reported for the selected pod.
  • Component categories and their Up, Degraded, Blocked, or Down counts.
  • Individual component status, severity, and detail text.

Use the breadcrumb links to return from a service to its module or to all modules.

Health status and service details for the controller pod

Deployment Activity

The Deployment Activity panel displays live license allocation and consumption. The overview shows the edition, number of licensed modules, license groups, and expiration warnings.

Select a group to review each capability's percentage used, used amount and allocation limit, expiration date when provided, and license edition. If consumption is unavailable, the card displays only the limit.

A group reports expiring soon when at least one capability expires within seven days. A capability expiration is critical at 30 days or less, warning at 90 days or less, and expired after its expiration date.

Use the Last polled value to confirm the freshness of live license data.

Viewing license usage for a module in Deployment Activity

Health Heatmap

The Health Heatmap displays as many as seven days of hourly module health. Each row represents one day, with the oldest day at the top.

Use the Module list to select one module or All. With All selected, each cell shows the worst available module status for that hour.

Unknown, null, and missing hourly values appear as No data.

Seven-day health heatmap for virtual appliance components.