3.4 Systems Manager OpsCenter & Incident Manager Runbooks

Key Takeaways

  • AWS Systems Manager OpsCenter centralizes operational and security issue management into unified OpsItems, pulling deep context from CloudWatch, Config, and CloudTrail.

  • OpsCenter mitigates alert fatigue through native deduplication, merging repeating or related operational alerts into a single active OpsItem using source and deduplication strings.

  • Systems Manager Incident Manager (closed to new customers since November 7, 2025) runs Response Plans that combine on-call contacts, escalation plans, chat channels, and automated SSM runbooks for existing customers.

  • Amazon Q Developer in chat applications (formerly Amazon Q Developer in chat applications) connects incident channels in Slack or Microsoft Teams to AWS and can run AWS CLI commands, limited by both the channel IAM role and channel guardrail policies.

  • Incident Manager Post-Incident Analysis (PIM) automatically aggregates chronological timelines of alarms, responder actions, and metrics, converting operational gaps into tracked remediation action items.

Last updated: September 2026

Systems Manager OpsCenter & Incident Manager Runbooks

When a major security incident strikes, technological containment is only half the battle. Responders must also coordinate cross-functional communications, page on-call subject matter experts, manage operational work items, track diagnostic telemetry, and maintain an immutable timeline of events for regulatory post-mortems. Without centralized incident management tooling, responders waste precious time switching between fragmented AWS consoles, misplacing critical context, and causing alert fatigue. AWS Systems Manager OpsCenter and AWS Systems Manager Incident Manager provide the native AWS control plane for operational issue tracking, automated paging escalation, collaborative war rooms, and post-incident analysis.


Important

Service status (2026). AWS stopped accepting new Incident Manager customers on November 7, 2025. Existing customers can keep using it, but it gets no new features. AWS's migration guidance points to OpsCenter for managing operational issues and to partner tools (PagerDuty, Jira Service Management, ServiceNow) for automated paging. AWS Chatbot is now Amazon Q Developer in chat applications. The exam guide names OpsCenter explicitly (skill 2.1.1), so know what each tool does, but design new paging workflows around OpsCenter plus a partner or EventBridge-driven integration.

OpsCenter: Centralized Operational Work Item Management

AWS Systems Manager OpsCenter provides a central operational work management console that aggregates, investigates, and resolves operational and security issues. Rather than requiring engineers to navigate between Amazon CloudWatch alarms, AWS Config non-compliance alerts, and AWS Security Hub findings, OpsCenter normalizes these events into unified work units called OpsItems.

+-------------------------------------------------------------+
| Operational Sources:                                        |
| • Amazon CloudWatch Alarms                                  |
| • AWS Security Hub Findings                                 |
| • AWS Config Non-Compliant Rules                            |
| • Amazon EventBridge Custom Rules                           |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| OpsCenter Deduplication Engine:                             |
| Evaluates source + dedup-string                             |
| • If matching open OpsItem exists -> Increment count        |
| • If no matching OpsItem exists   -> Create new OpsItem     |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| Unified OpsItem (Operational Context):                      |
| • Resource Metadata (AWS Config)                            |
| • CloudWatch Metric Graphs & Alarms                         |
| • CloudTrail API Logs                                       |
| • Associated SSM Automation Runbooks                        |
+-------------------------------------------------------------+

The Structure of an OpsItem

An OpsItem captures the complete operational state of an issue:

  • Title & Description: Clear summary of the operational disruption.
  • Source: The AWS service or custom system that generated the item (e.g., /aws/securityhub, /aws/cloudwatch).
  • Category & Severity: Classified into categories (such as Security, Availability, Performance, Cost) and given a priority from 1 (highest) to 5 plus a severity label such as Critical, High, Medium, or Low.
  • Related Resources: Amazon Resource Names (ARNs) of all impacted infrastructure components.
  • Operational Data: Arbitrary key-value pairs carrying payload diagnostics, finding attributes, or vulnerability scores.

Native Deduplication: Eliminating Alert Fatigue

A common failure in incident response is the "alert storm." When a core network component or database degrades, hundreds of downstream microservices emit cascading alarms. If each alarm generates a distinct ticket, responders become overwhelmed.

OpsCenter resolves this through built-in deduplication:

  • When an alert attempts to create an OpsItem, OpsCenter evaluates the alert's source and dedup-string.
  • If an active (status Open or In Progress) OpsItem already exists with that deduplication key, OpsCenter does not create a new OpsItem.
  • Instead, the new event is associated with the existing OpsItem (for OpsItems created from EventBridge rules, the dedup string travels in the /aws/dedup operational data key). This condenses thousands of repeating alerts into a single actionable ticket.

Deep Contextual Aggregation

When an engineer opens an OpsItem, OpsCenter automatically surfaces contextual diagnostic data without leaving the console:

  1. AWS Config History: Displays recent configuration changes to the affected resource, highlighting if an unauthorized security group modification occurred immediately before the alert.
  2. CloudWatch Metrics: Inlines performance and traffic graphs directly in the interface.
  3. CloudTrail Logs: Surfaces relevant API calls executed against the target resource.
  4. Associated Runbooks: Recommends specific Systems Manager Automation documents tailored to the OpsItem type (e.g., recommending AWS-RestartEC2Instance or AWSConfigRemediation-ConfigureS3BucketPublicAccessBlock).

Systems Manager Incident Manager: Crisis Response Orchestration

While OpsCenter manages operational issue backlogs and general work items, AWS Systems Manager Incident Manager is purpose-built for real-time, high-urgency crisis response and automated paging.

DimensionSystems Manager OpsCenterSystems Manager Incident Manager
Primary PurposeOperational work item tracking and issue managementReal-time crisis response, paging, and escalation
Urgency LevelAll severities (Sev 1 through Sev 5)Focused on high-impact disruptions (Sev 1 and Sev 2)
Key PrimitiveOpsItem (work tracking ticket)Incident & Response Plan (active crisis)
Paging & EscalationManual notificationAutomated multi-tier engagement plans (SMS, voice call)
CollaborationConsole notes and status updatesDedicated chat war rooms (Slack, Microsoft Teams)
Post-MortemStatus transition to ResolvedAutomated Post-Incident Analysis (PIM) with timeline

Response Plans: Declarative Crisis Playbooks

At the core of Incident Manager is the Response Plan. A response plan pre-defines the automated workflows executed the moment an incident is declared:

  • Incident Template: Pre-configures the default title, impact level (1-Critical, 2-High, 3-Medium, 4-Low, 5-Info), summary, and deduplication string.
  • Chat Channels: Binds the incident to dedicated Amazon Q Developer in chat applications collaboration channels.
  • Engagements: Specifies the on-call contacts and escalation plans to page.
  • Automated Runbooks: Designates Systems Manager Automation documents to execute immediately upon incident creation.

Example: Incident Manager Response Plan CLI Definition

{
  "name": "SecOpsCriticalContainmentPlan",
  "displayName": "SecOps Severity 1 Critical Containment Plan",
  "incidentTemplate": {
    "title": "Security Incident: Unauthorized Privilege Escalation",
    "impact": 1,
    "summary": "Critical privilege escalation detected. Initiating immediate responder paging and automated host snapshotting.",
    "dedupeString": "SecOps-PrivEscalation"
  },
  "chatChannel": {
    "chatbotSns": [
      "arn:aws:sns:us-east-1:111122223333:SecOpsChatbotIncidentBridgeTopic"
    ]
  },
  "engagements": [
    "arn:aws:ssm-contacts:us-east-1:111122223333:contact/secops-escalation-plan"
  ],
  "actions": [
    {
      "ssmAutomation": {
        "documentName": "AWS-CreateSnapshot",
        "roleArn": "arn:aws:iam::111122223333:role/IncidentManagerAutomationServiceRole",
        "documentVersion": "$DEFAULT"
      }
    }
  ]
}

Engagement Plans & Escalation Contacts

When a severity 1 incident occurs at 2:00 AM, broadcasting a generic email to an engineering team is ineffective. Incident Manager provides a structured, time-staged Engagement Plan and on-call rotation framework:

+-------------------------------------------------------------+
| Incident Declared (Severity 1)                              |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| Stage 1 (0 Minutes Elapsed):                                |
| Send immediate SMS & Mobile Push to Primary On-Call SecOps  |
+-------------------------------------------------------------+
                               |
             [Unacknowledged after 5 minutes]
                               v
+-------------------------------------------------------------+
| Stage 2 (5 Minutes Elapsed):                                |
| Initiate Automated Voice Phone Call to Secondary Responder  |
+-------------------------------------------------------------+
                               |
             [Unacknowledged after 15 minutes]
                               v
+-------------------------------------------------------------+
| Stage 3 (15 Minutes Elapsed):                               |
| Page Security Operations Director & Cloud Architecture Lead |
+-------------------------------------------------------------+

Contact Profiles & Communication Channels

Each responder profile in Incident Manager supports multiple communication channels:

  • SMS Text Messages: Delivers concise alert summaries.
  • Voice Phone Calls: Automated voice calls that read incident titles and prompt for tone acknowledgment.
  • Email Notifications: Delivers full incident context and links.
  • Mobile App Push Notifications: Alerts the AWS Console mobile application on responder smartphones.

Staged Escalation Schedules

Responders define tiered engagement plans with progressive wait durations:

  • Stage 1: Contacts the primary on-call engineer immediately via SMS and push notification.
  • Wait Interval: Incident Manager pauses for a configurable duration (e.g., 5 minutes) while waiting for an acknowledgment token.
  • Stage 2: If the primary responder does not acknowledge the page within 5 minutes, Incident Manager automatically initiates a voice call to the secondary engineer.
  • Cross-Team Escalation: If the incident remains uncontained after 15 minutes, the plan escalates ownership to Tier 2 Cloud Infrastructure Leads and Tier 3 Security Principals.

Collaborative Incident Triage with Amazon Q Developer in chat applications

During a crisis, incident responders must communicate in a shared, real-time war room rather than relying on disparate email threads. Incident Manager integrates natively with Amazon Q Developer in chat applications to bridge incident telemetry into Slack or Microsoft Teams.

Dedicated War Room Provisioning

When a Response Plan activates, Amazon Q Developer in chat applications automatically posts incident notifications into designated channels or creates an incident-specific war room channel. Channel members receive real-time updates as metrics spike, alarms transition, or runbooks complete.

Bi-Directional ChatOps Command Execution

Amazon Q Developer in chat applications is not merely a passive notification pipe; it supports bi-directional operational commands. Responders can query AWS infrastructure and execute remediation actions directly from chat:

  • Acknowledge pages: Responders acknowledge an engagement (by SMS reply, voice prompt, or in the console), which halts further escalation stages.
  • Inspect Infrastructure: Responders query live resource states (e.g., @aws ec2 describe-instances --filters ...).
  • Execute Runbooks: Responders trigger remediation runbooks without opening the AWS Management Console:
    @aws ssm start-automation-execution --document-name AWS-StopEC2Instance --parameters InstanceId=i-0123456789abcdef0
    

ChatOps Governance and Channel Guardrails

Executing operational commands from chat introduces security risks if unmonitored. Amazon Q Developer in chat applications enforces strict dual-layer authorization:

  1. Chatbot IAM Execution Role: Defines the maximum AWS permissions granted to the Chatbot service.
  2. Channel Guardrail Policies: An IAM policy applied directly to the chat channel. Even if an engineer has administrative privileges in their personal AWS account, any command executed within the chat channel is restricted by the channel guardrail. This prevents unauthorized users in the chat room from executing destructive commands (such as deleting databases or modifying IAM permissions).

Post-Incident Analysis (PIM / Post-Mortems) & Action Items

The ultimate goal of incident response is learning from failure. When an incident is resolved in Incident Manager, the service transitions to the Post-Incident Analysis (PIM) phase.

+-----------------------+      +-----------------------+      +-----------------------+
| 1. Auto-Timeline      | ---> | 2. Structured Review  | ---> | 3. Action Items       |
| • Alarms triggered    |      | • Root cause analysis |      | • Create OpsItems     |
| • Pages acknowledged  |      | • Detection lag       |      | • Sync with Jira /    |
| • Chatbot commands    |      | • Containment gaps    |      |   ServiceNow tickets  |
+-----------------------+      +-----------------------+      +-----------------------+

1. Automated Chronological Timeline Generation

Manually compiling a timeline after an incident is tedious and error-prone. Incident Manager automatically synthesizes an immutable chronological timeline of the incident by correlating:

  • The exact timestamp when the initiating CloudWatch alarm fired.
  • When the incident was declared and which response plan triggered.
  • When each responder was paged and when they acknowledged the notification.
  • Every command executed through Amazon Q Developer in chat applications during the triage.
  • Systems Manager Automation runbook start and completion timestamps.
  • Manual notes and diagnostic attachments added by responders.

2. Structured Post-Mortem Questionnaire

Incident Manager guides the team through a structured, blameless post-mortem framework:

  • Incident Summary: What occurred, and what was the customer impact?
  • Root Cause & Contributing Factors: What technical or procedural factors allowed the incident to occur?
  • Detection Metrics: What was the Mean Time to Detect (MTTD)? Did monitoring fire promptly?
  • Recovery Metrics: What was the Mean Time to Resolve (MTTR)? Did automated runbooks function as expected?
  • What Went Well / What Went Poorly: Qualitative assessment of team performance.

3. Action Item Tracking in OpsCenter

Lessons learned are useless if they remain buried in a post-mortem document. Incident Manager allows responders to convert post-incident recommendations directly into tracked action items:

  • Each action item generates a linked OpsItem in OpsCenter (or syncs bidirectionally with enterprise issue trackers like Jira or ServiceNow).
  • Action items track owners, due dates, and remediation progress, ensuring that architectural vulnerabilities and runbook gaps are closed before closing the review.

Investigation Runbooks in SageMaker AI Notebooks

Skill 2.1.1 also lists Amazon SageMaker AI notebooks as a way to build response runbooks. A Jupyter notebook is an executable playbook: each cell holds a documented step (an Athena query over CloudTrail or Security Lake tables, a boto3 call that lists a principal's recent AssumeRole events, a chart of VPC Flow Log bytes), and the output stays next to the step as a record of what the responder saw.

Design points that the exam cares about:

  • Run notebooks in the security tooling or forensics account, with an execution role that has read access to logs and only the containment permissions the playbook needs.
  • Put the notebook instance or SageMaker Studio domain in VPC-only mode with VPC endpoints, and encrypt its storage with a KMS key, because notebooks will hold sensitive evidence.
  • Keep playbooks in Git so they are reviewed and versioned like code, which also guards against the runbook drift covered in section 3.2.
  • Save notebook outputs to an Object Lock-protected evidence bucket when they become part of an investigation record.

Comparison of Incident Management & Tracking Tools

FeatureSystems Manager OpsCenterSystems Manager Incident ManagerAWS Health Dashboard
Primary RoleCentral operational work item management (OpsItems)Real-time crisis response, paging, and war roomsCloud service health and scheduled maintenance visibility
Core FocusOperational tickets, deduplication, diagnosticsHigh-severity incident management and escalationAWS infrastructure availability and event notifications
Paging / On-CallNo native paging (requires external tools)Native multi-tier paging (SMS, voice call, push)No on-call paging (emits EventBridge events)
Chat IntegrationNo direct chat bridgeNative chat integration (Slack, Microsoft Teams)Indirect via EventBridge to Amazon Q Developer in chat applications
Post-MortemBasic ticket resolutionAutomated Post-Incident Analysis with timelineService health event summary published by AWS
Automation LinkRecommends SSM Automation runbooksAutomatically executes SSM Automation runbooksCan trigger EventBridge remediation rules

Exam Tips & Common Traps

Important

Distinguish OpsCenter vs. Incident Manager: This is one of the most frequently tested distinctions in Domain 2. Remember:

  • Choose OpsCenter when the scenario requires centralizing operational work items (OpsItems), deduplicating repeating alerts from CloudWatch/Config/Security Hub, or viewing unified resource diagnostic data.
  • Choose Incident Manager when the scenario requires real-time on-call paging, staged escalation schedules (SMS to voice call), collaborative chat war rooms via Amazon Q Developer in chat applications, or automated post-incident analysis (post-mortems).

Warning

Common Trap on Amazon Q Developer in chat applications Permissions: Do not assume that any user in a Slack channel can execute arbitrary AWS commands if Amazon Q Developer in chat applications is configured. Amazon Q Developer in chat applications strictly evaluates Channel Guardrail Policies alongside its IAM execution role. If an action is not permitted by both the IAM execution role and the channel guardrail policy, the command is denied.

Tip

OpsCenter Deduplication Keys: When configuring custom EventBridge rules to create OpsItems, always provide a consistent dedup-string. Without a deduplication string, OpsCenter cannot correlate related alerts, leading to alert flooding during cascading system outages.

Loading diagram...
Systems Manager Incident Manager Operational Architecture
Test Your Knowledge

A security operations center (SOC) is overwhelmed by hundreds of alerts within fifteen minutes when a core microservice database fails, triggering cascading connection timeout alarms across dozens of dependent services. The SOC lead wants a centralized AWS service that aggregates these alerts into a single operational work item, deduplicates repeating alarms based on a common source key, and displays diagnostic CloudWatch metric graphs alongside recommended remediation runbooks. Which AWS service should the lead deploy?

A

AWS Systems Manager Incident Manager

B

AWS Trusted Advisor

C

AWS Security Hub

D

AWS Systems Manager OpsCenter

Test Your Knowledge

An enterprise that has used AWS Systems Manager Incident Manager in its security account since 2024 requires a structured escalation procedure for critical production security incidents. When a Severity 1 incident is declared, the primary security on-call engineer must be paged via SMS immediately. If the page is not acknowledged within five minutes, the system must automatically initiate a voice phone call to the secondary engineer; if still unacknowledged after fifteen minutes, it must page the Security Operations Director. Which AWS service and feature natively satisfies these staged escalation requirements?

A

Amazon SNS with delivery retry policies and dead-letter queues configured for each responder endpoint.

B

AWS Systems Manager Incident Manager with Engagement Plans and Escalation Contacts.

C

Amazon CloudWatch composite alarms configured with multiple Amazon SNS action endpoints.

D

AWS Step Functions Express Workflows invoking Amazon Connect automated contact flows.

Test Your Knowledge

During an active security incident, an engineering team uses a Slack war room connected to Amazon Q Developer in chat applications to collaborate on triage. An engineer attempts to execute a remediation runbook directly within the Slack channel by typing '@aws ssm start-automation-execution --document-name AWS-RestartEC2Instance'. The command fails with an 'Access Denied' authorization error, even though the engineer's personal IAM user has full administrative privileges in the AWS account. What is the most likely cause of this failure?

A

Amazon Q Developer in chat applications only supports read-only operations and cannot trigger Systems Manager Automation runbooks under any circumstances.

B

The engineer must be concurrently logged into the AWS Management Console in the same browser session where Slack is running.

C

The Amazon Q Developer in chat applications channel configuration uses an IAM execution role or channel guardrail policy that does not permit ssm:StartAutomationExecution.

D

Slack commands cannot invoke SSM documents unless the target EC2 instance is configured with an active public IP address.

Test Your Knowledge

Following the resolution of a major security disruption, a security architect must conduct a post-incident review (post-mortem). The architect needs to compile an accurate, chronological timeline of events—including when alarms fired, when responders acknowledged pages, which Chatbot commands were executed, and when automated runbooks completed—while tracking corrective action items assigned to engineering teams. Which tool provides this automated timeline generation and post-incident review workflow natively?

A

AWS Systems Manager Incident Manager Post-Incident Analysis

B

AWS CloudTrail Lake

C

Amazon Detective

D

AWS Systems Manager Inventory

Sections you finish are checked off in the contents.