Build operational confidence
Use operational scenarios and team capability checks to establish confidence in Splunk Observability Cloud before cutover.
Overview
Prepare operations and on-call teams to use Splunk Observability Cloud as the primary observability platform.
Prerequisites
Confirm that:
- Complete controlled component testing.
- Activate or complete the dual-platform validation period.
- Operations personnel have the required Splunk Observability Cloud permissions.
- Ensure access to critical dashboards, detectors, notification routes, and runbooks.
- Provide application and platform specialist support for practice sessions.
Use Splunk Observability Cloud as the primary investigation tool
During the shadow period:
- Begin each investigation in Splunk Observability Cloud.
- Use Splunk AppDynamics only to compare results or confirm an identified migration gap.
- Record any missing data, permission, navigation guidance, or monitoring capability.
- Update the applicable configuration or runbook.
- Repeat the workflow until the team can complete it using the target platform.
Practice priority operational scenarios
Exercise representative scenarios for the migration wave:
| Scenario | Expected workflow |
|---|---|
| High service latency | Start from the detector or dashboard, identify the affected service or Business Transaction, inspect traces, and locate the contributing span or dependency |
| Error-rate increase | Find error traces, review exception details and affected endpoints, and identify the likely source |
| Dependency failure | Use Service Map and traces to identify the upstream and downstream impact |
| Infrastructure-related service issue | Correlate service behavior with available host, container, or Kubernetes context |
| Alert triage | Receive the notification, open the detector, assess severity and scope, follow the runbook, and escalate when required |
| Missing telemetry | Check the service time range, trace and metric availability, Collector or agent status, and recent deployment changes |
Include other application-specific scenarios from existing incident records and runbooks.
Confirm team capabilities
Confirm that the responsible teams can:
- Find services, Business Transactions, endpoints, traces, dashboards, detectors, and alerts.
- Filter by the approved environment, namespace, service, workflow, and instance dimensions.
- Investigate latency, errors, dependencies, and missing telemetry.
- Interpret detector conditions, severities, and clear behavior.
- Create or safely tune an approved detector or dashboard.
- Verify notification routing and escalation.
- Use Tag Spotlight, Service Map, Trace Analyzer, and supported profiling features when applicable.
- Contact the correct application owner, platform owner, or Splunk Support contact.
Update operational documentation
Update:
- Incident-response runbooks.
- Alert ownership and escalation paths.
- Dashboard and detector ownership.
- Notification-routing documentation.
- Known differences and accepted gaps.
- Rollback and cutover contacts.
- Support escalation information.
Do not treat attendance at a training session as operational readiness. Require team members to complete the priority workflows using representative telemetry.
Completion criteria
Establish operational confidence when:
- Each responsible team completes the required scenarios in Splunk Observability Cloud.
- Teams no longer require undocumented assistance for priority workflows.
- Resolve or assign missing permissions and gaps.
- Runbooks and escalation procedures reference the target platform.
- Operations and on-call owners approve readiness for cutover.