Build operational confidence

Use operational scenarios and team capability checks to establish confidence in Splunk Observability Cloud before cutover.

Overview

Prepare operations and on-call teams to use Splunk Observability Cloud as the primary observability platform.

Prerequisites

Confirm that:

  • Complete controlled component testing.
  • Activate or complete the dual-platform validation period.
  • Operations personnel have the required Splunk Observability Cloud permissions.
  • Ensure access to critical dashboards, detectors, notification routes, and runbooks.
  • Provide application and platform specialist support for practice sessions.

Use Splunk Observability Cloud as the primary investigation tool

During the shadow period:

  1. Begin each investigation in Splunk Observability Cloud.
  2. Use Splunk AppDynamics only to compare results or confirm an identified migration gap.
  3. Record any missing data, permission, navigation guidance, or monitoring capability.
  4. Update the applicable configuration or runbook.
  5. Repeat the workflow until the team can complete it using the target platform.

Practice priority operational scenarios

Exercise representative scenarios for the migration wave:

Scenario Expected workflow
High service latency Start from the detector or dashboard, identify the affected service or Business Transaction, inspect traces, and locate the contributing span or dependency
Error-rate increase Find error traces, review exception details and affected endpoints, and identify the likely source
Dependency failure Use Service Map and traces to identify the upstream and downstream impact
Infrastructure-related service issue Correlate service behavior with available host, container, or Kubernetes context
Alert triage Receive the notification, open the detector, assess severity and scope, follow the runbook, and escalate when required
Missing telemetry Check the service time range, trace and metric availability, Collector or agent status, and recent deployment changes

Include other application-specific scenarios from existing incident records and runbooks.

Confirm team capabilities

Confirm that the responsible teams can:

  • Find services, Business Transactions, endpoints, traces, dashboards, detectors, and alerts.
  • Filter by the approved environment, namespace, service, workflow, and instance dimensions.
  • Investigate latency, errors, dependencies, and missing telemetry.
  • Interpret detector conditions, severities, and clear behavior.
  • Create or safely tune an approved detector or dashboard.
  • Verify notification routing and escalation.
  • Use Tag Spotlight, Service Map, Trace Analyzer, and supported profiling features when applicable.
  • Contact the correct application owner, platform owner, or Splunk Support contact.

Update operational documentation

Update:

  • Incident-response runbooks.
  • Alert ownership and escalation paths.
  • Dashboard and detector ownership.
  • Notification-routing documentation.
  • Known differences and accepted gaps.
  • Rollback and cutover contacts.
  • Support escalation information.

Do not treat attendance at a training session as operational readiness. Require team members to complete the priority workflows using representative telemetry.

Completion criteria

Establish operational confidence when:

  • Each responsible team completes the required scenarios in Splunk Observability Cloud.
  • Teams no longer require undocumented assistance for priority workflows.
  • Resolve or assign missing permissions and gaps.
  • Runbooks and escalation procedures reference the target platform.
  • Operations and on-call owners approve readiness for cutover.