3.3 Automated Incident Remediation with EventBridge & Step Functions

Key Takeaways

  • Event-driven remediation matches high-severity GuardDuty findings and Security Hub CSPM findings (ASFF) with Amazon EventBridge rules to achieve sub-minute automated response.

  • AWS Step Functions state machines coordinate multi-step, distributed remediation workflows with visual branching logic, deterministic retry policies, exponential backoff, and error-handling Catch blocks.

  • Systems Manager Automation runbooks such as AWS-QuarantineEC2Instance, AWSSupport-ContainEC2Instance, AWS-CreateSnapshot, and AWSConfigRemediation-ConfigureS3BucketPublicAccessBlock are auditable remediation building blocks that run across accounts.

  • AWS Lambda runs targeted response logic, such as revoking an IAM role's active sessions with an inline deny conditioned on aws:TokenIssueTime being earlier than the revocation time.

  • Enterprise remediation workflows must incorporate safety guardrails: dry-run mode verification, idempotency, automated rollback routines, and human-in-the-loop approval gates using Step Functions .waitForTaskToken callback patterns.

Last updated: September 2026

Automated Incident Remediation with EventBridge & Step Functions

When an adversary breaches a cloud workload, automated script execution operates in seconds, while human security analysts require minutes or hours to detect, investigate, and manually contain the threat. This disparity in operational velocity is known as adversary dwell time. To eliminate dwell time and mitigate potential damage, cloud security architectures must incorporate event-driven automated remediation. By coupling real-time security telemetry with serverless orchestration engines, organizations can detect an intrusion, quarantine compromised resources, revoke leaked credentials, and capture forensic evidence in sub-minute timeframes.


Event-Driven Security Remediation Architecture

In AWS, automated security remediation follows an event-driven architecture centered on Amazon EventBridge, which acts as the universal serverless event bus:

+-------------------------------------------------------------+
| 1. Threat Detection Engine:                                 |
|    • Amazon GuardDuty (e.g., UnauthorizedAccess:EC2)        |
|    • AWS Security Hub (ASFF finding: Compliance FAILED)     |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| 2. Event Routing (Amazon EventBridge):                      |
|    Matches JSON Event Pattern (Severity >= 7.0)             |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| 3. Workflow Orchestration (AWS Step Functions):             |
|    Evaluates Resource Type -> Enforces Safety Guardrails    |
+-------------------------------------------------------------+
                               |
        +----------------------+----------------------+
        |                                             |
        v                                             v
+-------------------------------+   +---------------------------------+
| 4a. AWS SSM Automation:       |   | 4b. AWS Lambda:                 |
| • AWS-StopEC2Instance         |   | • Revoke Active STS Sessions    |
| • AWS-CreateSnapshot          |   | • Disable IAM Access Keys       |
| • S3 Block Public Access      |   | • Quarantine Network Interface  |
+-------------------------------+   +---------------------------------+

Ingesting Normalized Findings: The AWS Security Finding Format (ASFF)

Different security services historically produced divergent alert formats. AWS Security Hub resolves this by normalizing all findings—whether originating from native services like GuardDuty, Macie, and Inspector, or from third-party firewalls—into the standardized AWS Security Finding Format (ASFF). Key ASFF JSON fields evaluated by automated remediation rules include:

  • Id: The unique identifier for the specific finding.
  • ProductArn: Identifies the emitting tool (e.g., arn:aws:securityhub:us-east-1::product/aws/guardduty).
  • AwsAccountId: The account where the incident occurred.
  • Types: An array classifying the threat (e.g., ["TTPs/Initial Access/UnauthorizedAccess:EC2-TorIP"]).
  • Severity.Label: The normalized severity string (INFORMATIONAL, LOW, MEDIUM, HIGH, CRITICAL).
  • Severity.Normalized: A numerical score from 0 to 100, where scores 70.0 to 89.9 represent HIGH and 90.0 to 100.0 represent CRITICAL.
  • Resources: Detailed metadata describing affected EC2 instances, S3 buckets, IAM roles, or Lambda functions.
  • Compliance.Status: In compliance checks, indicates PASSED, WARNING, FAILED, or NOT_AVAILABLE.

Constructing EventBridge Rule Event Patterns

Amazon EventBridge evaluates incoming JSON events against declarative filter patterns. To trigger remediation only for high-priority threats and avoid flooding automation runners with informational noise, security teams construct specific event patterns:

Example: EventBridge Pattern Matching High-Severity GuardDuty Findings

{
  "source": ["aws.guardduty"],
  "detail-type": ["GuardDuty Finding"],
  "detail": {
    "severity": [
      { "numeric": [">=", 7.0] }
    ],
    "type": [
      { "prefix": "UnauthorizedAccess:EC2" },
      { "prefix": "Recon:IAMUser" },
      { "prefix": "CryptoCurrency:EC2" }
    ]
  }
}

Example: EventBridge Pattern Matching Failed Security Hub S3 Compliance Findings

{
  "source": ["aws.securityhub"],
  "detail-type": ["Security Hub Findings - Imported"],
  "detail": {
    "findings": {
      "Compliance": {
        "Status": ["FAILED"]
      },
      "GeneratorId": [
        { "prefix": "aws-foundational-security-best-practices/v/1.0.0/S3." }
      ],
      "Severity": {
        "Label": ["CRITICAL", "HIGH"]
      }
    }
  }
}

Cross-Account Event Aggregation

In multi-account environments, findings emitted in spoke workload accounts should not trigger uncoordinated local remediation. Instead, spoke accounts configure EventBridge rules that forward all security events to a centralized Security Event Bus hosted in the Security Tooling account. Remediation workflows execute centrally, maintaining unified audit logging and centralized control.


AWS Step Functions: Multi-Step Remediation Orchestration

While a single AWS Lambda function can execute a basic script, real-world incident containment requires multi-step, sequential, and parallel coordination across multiple systems. Relying on a single monolithic Lambda script introduces severe fragility: if one API call times out or fails, the entire script may crash mid-execution, leaving resources partially modified in an undefined state.

AWS Step Functions state machines provide the ideal enterprise remediation orchestrator:

  • Visual Workflow Management: Every state transition, execution path, and error is recorded graphically in real time.
  • Built-in Error Handling & Retries: Native Retry blocks support exponential backoff, jitter, and error-specific handling without writing custom code.
  • Catch Blocks & Fallback States: If a primary remediation step fails, a Catch block routes execution to an alert state that pages the on-call engineer.
  • Choice States: The state machine dynamically evaluates the finding metadata to branch between production and non-production accounts or between compute, storage, and identity remediation paths.

Standard vs. Express Workflows for Incident Response

Step Functions provides two distinct workflow types:

FeatureStandard WorkflowsExpress Workflows
Execution SemanticsExactly-once executionAt-least-once execution
Maximum DurationUp to 1 yearUp to 5 minutes
Execution HistoryFull visual history inspected via console/API for 90 daysExecution logs pushed to CloudWatch Logs
Human Approval SupportSupported via .waitForTaskToken callback patternNot supported
Throughput / PricingBilled per state transition; moderate throughputBilled per GB-second; ultra-high throughput (100k+/sec)
Primary IR Use CaseComplex multi-account containment, approval gates, forensicsHigh-frequency log scrubbing, rapid alert triage

For security remediation involving human approval gates, EBS snapshot creation, and deep forensic capture, Standard Workflows are required.


AWS Systems Manager Automation Documents

AWS Systems Manager (SSM) Automation allows engineers to safely execute standardized operational and security runbooks at scale. Instead of reinventing custom remediation scripts, security teams leverage pre-built, AWS-managed SSM documents:

1. AWSConfigRemediation-ConfigureS3BucketPublicAccessBlock

When an S3 bucket is inadvertently made public, this runbook turns on the bucket-level S3 Block Public Access settings (BlockPublicAcls, IgnorePublicAcls, BlockPublicPolicy, RestrictPublicBuckets). The older AWS-DisableS3BucketPublicReadWrite runbook serves the same purpose for a single bucket.

2. AWS-StopEC2Instance

Gracefully stops a target EC2 instance. In an automated containment workflow, this action is typically sequenced after EBS snapshot creation and volatile memory capture to freeze persistent block storage.

3. AWS-CreateSnapshot

Creates a point-in-time snapshot of one EBS volume (you pass the volume ID). Loop over each attached volume, or call the EC2 CreateSnapshots API to capture all of an instance's volumes at one crash-consistent point. These snapshots are tagged with incident identifiers and can be automatically shared with the isolated Forensics account.

4. AWS-QuarantineEC2Instance and AWSSupport-ContainEC2Instance

AWS-QuarantineEC2Instance assigns a security group that allows no inbound or outbound traffic; AWSSupport-ContainEC2Instance is a broader containment runbook for an instance. Detaching the IAM instance profile (the EC2 DisassociateIamInstanceProfile API, typically called from Lambda or a custom Automation step) stops IMDS from serving new role credentials, but credentials already stolen stay valid until you revoke the role's active sessions.

Cross-Account SSM Automation Execution

To execute SSM runbooks across member accounts from the centralized Security Tooling account:

  1. The centralized Step Functions state machine calls the SSM StartAutomationExecution API.
  2. The call specifies the AutomationAssumeRole parameter pointing to a pre-provisioned role in the spoke account (e.g., arn:aws:iam::444455556666:role/SSMAutomationSpokeExecutionRole).
  3. The spoke role executes the document locally, ensuring the security account does not require broad direct permissions inside the spoke account.

AWS Lambda for High-Speed Targeted Remediation

For microsecond-level execution speeds and custom containment logic, AWS Lambda complements SSM Automation. A prime example is instantly revoking compromised IAM user credentials.

The Critical Nuance of Active STS Session Invalidation

When an IAM user's long-term access keys are leaked, junior engineers often make the mistake of simply calling iam:DeleteAccessKey or changing the user's password. This does not terminate an active attack.

If the adversary used the compromised access key to generate temporary security credentials via AWS STS (such as calling sts:GetSessionToken or sts:AssumeRole), those temporary session tokens remain valid until their expiration time (which can be up to 12 or 36 hours). The attacker continues executing API calls unimpeded.

To immediately and completely invalidate all existing temporary STS sessions, the remediation Lambda function must attach an inline session revocation policy to the user or role. The policy uses an explicit Deny with an aws:TokenIssueTime condition:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ImmediatelyRevokeAllActiveSessions",
      "Effect": "Deny",
      "Action": "*",
      "Resource": "*",
      "Condition": {
        "DateLessThan": {
          "aws:TokenIssueTime": "2026-09-29T12:00:00Z"
        }
      }
    }
  ]
}

How it works: aws:TokenIssueTime is the time the temporary credential was issued. Any request signed with a credential issued before 2026-09-29T12:00:00Z satisfies DateLessThan, so the explicit Deny blocks it, while credentials issued afterward work normally. Do not use aws:CurrentTime here: that key is the time of the request, so the deny would block everyone until the timestamp passes and then stop protecting you. Requests signed directly with long-term access keys carry no aws:TokenIssueTime, so deactivate or delete those keys as well. For roles, the IAM console's Revoke active sessions action attaches this exact pattern as the AWSRevokeOlderSessions inline policy.


Safety Guardrails, Approval Gates & Rollback Mechanisms

Uncontrolled automated remediation is dangerous. If an automated script encounters a false-positive alert and aggressively shuts down a core production database or revokes the credentials of an enterprise application role, the remediation script causes more damage than the threat it was designed to contain. Enterprise remediation workflows must incorporate strict safety guardrails:

+-------------------------------------------------------------+
| EventBridge Trigger: High-Severity Finding Received         |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| Step Functions Choice State: Is this a Production Workload? |
+-------------------------------------------------------------+
              |                                 |
     [YES]    |                                 |   [NO / Staging]
              v                                 v
+---------------------------+     +---------------------------+
| Human Approval Gate:      |     | Immediate Execution:      |
| • Pause via Task Token    |     | • Isolate Instance        |
| • Send Email / Slack Link |     | • Revoke Credentials      |
+---------------------------+     +---------------------------+
              |                                 |
     [Approved by SecOps]                       |
              v                                 |
+-----------------------------------------------+             |
| Execute Containment with Dry-Run Validation   | <-----------+
+-----------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| Verify Finding Status -> Update Security Hub to RESOLVED    |
+-------------------------------------------------------------+

1. Human-in-the-Loop Approval Gates (.waitForTaskToken)

For production environments or destructive remediation actions (such as terminating compute or modifying public databases), Step Functions provides the Task Token Callback pattern:

  1. The state machine reaches a Task state that publishes an approval request to an Amazon SNS topic or Slack channel using arn:aws:states:::sns:publish.waitForTaskToken.
  2. Step Functions pauses execution and generates a unique, cryptographically signed TaskToken.
  3. The notification contains two URLs (e.g., pointing to an Amazon API Gateway endpoint with query parameters): one for Approve and one for Reject, embedding the task token.
  4. When a verified security engineer clicks "Approve", API Gateway triggers a Lambda function that calls sfn:SendTaskSuccess(taskToken). If clicked "Reject", it calls sfn:SendTaskFailure(taskToken).
  5. Upon receiving SendTaskSuccess, Step Functions unpauses and proceeds with the automated remediation steps.

2. Idempotency and Dry-Run Verification

Automated remediation scripts must be idempotent: executing the script multiple times against the same resource must produce the exact same final state without generating errors or corrupting infrastructure. Furthermore, scripts interacting with EC2 should execute API calls with --dry-run to verify that IAM execution roles have the necessary permissions and that target resource IDs exist before executing state alterations.

3. Automated Rollback and Finding Closure

If a containment step fails midway through execution, Step Functions Catch blocks trigger rollback actions (such as restoring original security groups or alerting the primary on-call engineer through the paging tool). Upon successful containment, the workflow calls securityhub:BatchUpdateFindings to update the finding's Workflow.Status to RESOLVED, documenting the automated remediation audit trail in Security Hub.


Restoring Availability with Amazon Application Recovery Controller (ARC)

Skill 2.1.4 also names Amazon Application Recovery Controller (ARC), because recovery is part of remediation. ARC provides:

  • Zonal shift: temporarily moves traffic for a supported resource (for example, an Application Load Balancer or Network Load Balancer with cross-zone settings) away from one Availability Zone. A shift is manual, temporary, and carries an expiration of up to three days that you can extend. Zonal autoshift lets AWS start the shift automatically when it detects an AZ impairment.
  • Routing control and Region switch: multi-Region recovery controls that move traffic between Regional replicas of an application, backed by readiness checks that confirm the standby is actually able to take load.

In an incident, zonal shift is a fast containment-plus-recovery lever: if every compromised instance sits in one AZ, shifting traffic away keeps customers served while responders snapshot and rebuild that capacity. Automate it the same way as other remediations: an EventBridge rule or Step Functions state calls the ARC API, and CloudTrail records who started the shift.

Comparison of Incident Remediation Technologies

FeatureAWS Systems Manager AutomationAWS Step FunctionsAWS Lambda
Primary StrengthPre-built, AWS-managed operational runbooksMulti-step orchestration, visual branching, human approvalsCustom, high-speed microsecond containment logic
Execution ModelDocument-based declarative runbookState machine workflow (Standard or Express)Serverless event-driven code execution
Max DurationLong-running; steps can pause for approvalUp to 1 year (Standard) / 5 minutes (Express)15 minutes max per execution
Human ApprovalSupported with the aws:approve actionSupported natively via .waitForTaskTokenMust be implemented manually via custom databases
Multi-AccountNative via AutomationAssumeRoleOrchestrates multi-account calls via SDK integrationRequires cross-account STS assume role in code
Best Used ForStandard instance/S3 tasks (e.g. S3 Block Public Access)End-to-end incident response pipelinesImmediate session revocation, custom API calls

Exam Tips & Common Traps

Important

Session Token Revocation Mechanics: The SCS-C03 exam regularly tests how to neutralize compromised IAM credentials. Remember: Deleting an IAM access key or changing a console password does NOT revoke active STS session tokens. If the question asks for immediate invalidation of all temporary security credentials, the correct answer must include attaching an explicit deny policy with an aws:TokenIssueTime condition (what the console's Revoke active sessions action does for roles), plus deactivating any leaked long-term keys.

Warning

Common Trap on Step Functions Workflow Types: If an exam scenario requires an automated remediation pipeline that includes human-in-the-loop approvals (such as requiring a manager to approve deleting a resource or isolating a production server) or tasks that wait for extended periods, do not select Express Workflows. Express Workflows have a 5-minute execution limit and do NOT support the .waitForTaskToken callback pattern. You must select Standard Workflows.

Tip

ASFF Finding Severity Enums: In EventBridge pattern matching, remember that AWS Security Hub findings use capitalized strings for the Severity.Label field (INFORMATIONAL, LOW, MEDIUM, HIGH, CRITICAL), whereas raw Amazon GuardDuty findings use numerical floats (severity: [{ numeric: [">=", 7.0] }]). Ensure your rule matches the schema of the emitting service.

Loading diagram...
End-to-End Automated Security Remediation Workflow
Test Your Knowledge

Amazon GuardDuty generates a high-severity finding indicating that an IAM user's long-term access keys have been compromised and are actively being used to make unauthorized API calls from a known malicious IP address (UnauthorizedAccess:IAMUser/MaliciousIPCaller). An automated remediation Lambda function triggers immediately. Which action must the Lambda function take to immediately prevent the attacker from making any further API calls using existing temporary security credentials previously issued to that user?

A

Call iam:DeleteAccessKey to delete the compromised access key and execute iam:UpdateLoginProfile to reset the user's console password.

B

Attach an AWS Organizations Service Control Policy (SCP) to the member account that contains an explicit deny for all IAM actions across the account.

C

Attach an inline IAM policy to the compromised user that explicitly denies all actions when aws:TokenIssueTime is earlier than the revocation timestamp, and deactivate the leaked access key.

D

Invoke the sts:RevokeToken API call specifying the access key ID to invalidate all active session tokens immediately.

Test Your Knowledge

A financial enterprise requires automated remediation for non-compliant, publicly accessible S3 buckets detected by AWS Security Hub. In staging accounts, public access must be blocked immediately without human intervention. In production accounts, the remediation workflow must pause, send an approval link to the on-call security manager via email, and only proceed to block public access after the manager clicks 'Approve'. Which AWS Step Functions design fulfills these requirements?

A

Deploy a Standard Workflow that evaluates the account environment tag; for production accounts, invoke an Amazon SNS task using the .waitForTaskToken pattern to pause execution until an API Gateway endpoint receives approval and calls SendTaskSuccess.

B

Deploy an Express Workflow with a Wait state of 24 hours that periodically polls an Amazon DynamoDB table where the security manager writes an approval flag.

C

Deploy two separate Lambda functions connected via Amazon SQS; the first function sends an email and sleeps for 30 minutes before the second function executes S3 Block Public Access.

D

Deploy an AWS Systems Manager Automation document that includes a manual approval parameter and pause the Amazon EventBridge rule directly in the source account.

Test Your Knowledge

A security engineer designs an Amazon EventBridge rule to route high-severity compliance findings from AWS Security Hub to an automated Step Functions remediation state machine. Which EventBridge event pattern correctly matches only findings with a severity label of HIGH or CRITICAL in the AWS Security Finding Format (ASFF)?

A

{"source": ["aws.guardduty"], "detail-type": ["Security Hub Finding"], "detail": {"severity": ["HIGH", "CRITICAL"]}}

B

{"source": ["aws.securityhub"], "detail-type": ["Security Alert"], "detail": {"findings": {"Severity": {"Label": ["High", "Critical"]}}}}

C

{"source": ["aws.securityhub"], "detail": {"status": ["FAILED"], "severity": [">=70"]}}

D

{"source": ["aws.securityhub"], "detail-type": ["Security Hub Findings - Imported"], "detail": {"findings": {"Severity": {"Label": ["HIGH", "CRITICAL"]}}}}

Test Your Knowledge

A global enterprise needs to standardize its EC2 incident containment procedures across 60 AWS member accounts. When an EC2 instance is identified as compromised, the containment runbook must detach the instance profile, attach a zero-traffic quarantine security group, and snapshot all attached EBS volumes. Which AWS service provides pre-built, managed automation runbooks for these exact tasks that can be executed centrally across member accounts?

A

AWS CloudFormation StackSets

B

AWS Systems Manager Automation

C

AWS OpsWorks Stacks

D

AWS Elastic Beanstalk

Sections you finish are checked off in the contents.