3.1 Incident Response Planning & Blast Radius Mitigation

Key Takeaways

  • Cloud incident response follows the NIST SP 800-61 lifecycle (Preparation; Detection and Analysis; Containment, Eradication, and Recovery; Post-Incident Activity); Rev. 3 (April 2025) maps the same work to the NIST CSF 2.0 functions.

  • Dedicated Security and Forensics accounts in AWS Organizations isolate the blast radius, preventing compromised workload credentials from tampering with audit logs or running forensic analysis.

  • Emergency break-glass IAM roles must be pre-provisioned in spoke accounts with strict trust policies that enforce MFA (aws:MultiFactorAuthPresent), limit session duration, and alert the SOC immediately upon assumption.

  • Blast radius containment combines Service Control Policies (SCPs) to protect defensive controls, isolated forensic VPCs without Internet Gateways, and Transit Gateway route table quarantine to prevent lateral movement.

  • Shield Advanced proactive engagement (Business or Enterprise Support, Route 53 health checks, at least one contact) lets the Shield Response Team (SRT) contact you during detected events; a role trusting drt.shield.amazonaws.com with AWSShieldDRTAccessPolicy lets it adjust AWS WAF rules.

Last updated: September 2026

Incident Response Planning & Blast Radius Mitigation

Cloud incident response differs fundamentally from traditional on-premises security operations. In an on-premises data center, incident response teams physically access hardware, tap network switches, or pull power cables to isolate compromised servers. In Amazon Web Services (AWS), the physical layer is managed entirely by AWS under the Shared Responsibility Model, while customer incident response operates programmatically through software-defined application programming interfaces (APIs), identity boundaries, and network virtualization controls. Preparing for security incidents in AWS requires architecting isolation boundaries, pre-authorizing emergency access channels, and establishing automated blast radius mitigation before an intrusion occurs.


The Cloud Incident Response Lifecycle (AWS CIRT Framework)

AWS incident response guidance builds on the NIST SP 800-61 incident handling lifecycle. Rev. 2 describes four phases (Preparation; Detection and Analysis; Containment, Eradication, and Recovery; Post-Incident Activity). Rev. 3, published in April 2025, reorganizes the same work around the NIST CSF 2.0 functions. The diagram below splits containment from eradication and recovery for teaching purposes:

+---------------+      +----------------------+      +---------------+
|  Preparation  | ---> | Detection & Analysis | ---> |  Containment  |
+---------------+      +----------------------+      +---------------+
                                                             |
                                                             v
+-----------------------+      +-------------------------------------+
| Post-Incident Activity | <--- |       Eradication & Recovery        |
+-----------------------+      +-------------------------------------+

1. Preparation

Preparation forms the bedrock of cloud incident readiness. Organizations must pre-provision identity roles, configure immutable logging pipelines, and ensure security responders have pre-authorized operational mechanisms across all AWS accounts. Key preparation controls include:

  • Deploying centralized AWS CloudTrail organization trails with log file integrity validation enabled and sending digests to an isolated S3 bucket protected by S3 Object Lock in Compliance mode.
  • Enabling Amazon GuardDuty, AWS Security Hub, and AWS Config across all accounts and AWS Regions using AWS Organizations delegated administrator accounts.
  • Pre-provisioning emergency cross-account IAM roles (break-glass roles) with appropriate trust policies and permissions boundaries.
  • Establishing documented incident playbooks and automated runbooks for common attack vectors, including credential compromise, ransomware encryption, and cryptomining compute abuse.

2. Detection & Analysis

Detection involves ingesting telemetry to identify deviations from normal baseline behavior, triaging alerts, and determining the scope of an incident. In AWS, this relies heavily on managed threat detection:

  • Amazon GuardDuty analyzes CloudTrail management events, VPC Flow Logs, and DNS query logs, plus S3 data events, EKS audit logs, Lambda network activity, and runtime telemetry when those protection plans are enabled, using machine learning and threat intelligence to identify compromised resources.
  • AWS Security Hub normalizes findings across AWS native security services and third-party tools into the AWS Security Finding Format (ASFF), calculating composite compliance scores and routing alerts.
  • Responders examine CloudTrail events to determine the compromised identity (userIdentity), source IP address (sourceIPAddress), user agent, requested API actions, and geographic anomalies.

3. Containment

Containment prevents the adversary from causing further damage, exfiltrating sensitive data, or pivoting laterally to other cloud resources. Containment is executed in two stages:

  • Short-Term Containment: Rapidly severing malicious communication channels. Examples include applying a restrictive quarantine Security Group to an EC2 instance, attaching an inline deny policy to a compromised IAM user or role, or invalidating active session tokens via AWS Security Token Service (STS).
  • Long-Term Containment: Isolating affected environments while preserving evidence. This includes detaching network interfaces, cloning block storage volumes for out-of-band analysis, and updating routing tables to route traffic away from compromised workloads.

4. Eradication & Recovery

Eradication removes all traces of the threat from the environment, while recovery safely restores clean business workloads to full operational status:

  • Terminating compromised virtual machines and replacing them with verified clean instances launched from trusted Golden AMIs through automated CI/CD deployment pipelines.
  • Revoking compromised long-term IAM access keys, rotating database master passwords stored in AWS Secrets Manager, and re-encrypting impacted data using new AWS Key Management Service (AWS KMS) customer managed keys (CMKs).
  • Restoring databases and file systems from verified, point-in-time snapshots stored in a dedicated AWS Backup vault protected by Backup Vault Lock.

5. Post-Incident Activity

Often called the "post-mortem" or "lessons learned" phase, this activity ensures that every incident drives lasting security improvement:

  • Conducting a blameless post-incident review involving engineering, security, and executive stakeholders.
  • Identifying detection gaps, procedural bottlenecks, or missing runbook automations.
  • Feeding corrective action items into engineering backlogs and updating incident response playbooks to prevent recurrence.

Multi-Account Architecture: Security and Forensics Accounts

A critical tenet of AWS security is using multiple AWS accounts to establish strict administrative, cryptographic, and network isolation boundaries. Operating workloads inside a single AWS account creates a monolithic blast radius: an attacker who achieves administrative access (AdministratorAccess) in that account can tamper with CloudTrail trails, delete backup snapshots, disable security monitoring, and destroy evidence.

Under AWS Organizations, enterprises organize accounts into a structured hierarchy of Organizational Units (OUs). Within the Security OU, two specialized accounts must be provisioned:

                    AWS Organizations (Root)
                               |
         +---------------------+---------------------+
         |                                           |
    Core OU (Security)                         Workloads OU
         |                                           |
    +----+----+----+                       +---------+---------+
    |         |    |                       |                   |
Security   Log  Forensics              Production         Development
Tooling  Archive Account                 Account            Account
Account  Account

The Log Archive Account

The Log Archive account acts as a tamper-resistant vault for all organization telemetry. Member accounts stream CloudTrail logs, VPC Flow Logs, and Route 53 query logs directly into central S3 buckets hosted in this account.

  • S3 Object Lock (Compliance Mode): Ensures that no user—including the AWS account root user or organization administrators—can delete or overwrite log files during the mandatory retention window.
  • KMS Customer Managed Key Isolation: The KMS key used to encrypt logs resides in the Log Archive account. Its key policy allows spoke accounts to perform kms:GenerateDataKey* operations but restricts kms:Decrypt and administrative permissions strictly to authorized log analysis roles.
  • MFA Delete: Configured on the S3 bucket versioning configuration to mandate hardware token authentication before deleting any object version.

The Security Tooling Account

The Security Tooling account serves as the operational headquarters for the Security Operations Center (SOC). It hosts centralized security services configured as Delegated Administrators in AWS Organizations:

  • AWS Security Hub Delegated Administrator: Aggregates findings from every member account and Region across the enterprise into a centralized dashboard.
  • Amazon GuardDuty Delegated Administrator: Centrally manages detector configurations, suppression rules, and malware protection across all member accounts.
  • Amazon Detective Delegated Administrator: Ingests telemetry across accounts to automatically construct graph-based security investigation models.

The Dedicated Forensics Account

The Forensics account is an air-gapped, isolated environment purpose-built for deep artifact analysis. Responders never perform forensic investigations inside the compromised production account because the adversary may monitor API calls, execute anti-forensic wiper scripts, or exploit local IAM permissions to tamper with memory dumps.

  • Isolated Forensics VPC: The forensics VPC is constructed without an Internet Gateway (IGW), without a NAT Gateway, and without Virtual Private Gateway connections to the corporate WAN. This guarantees zero outbound internet egress, completely eliminating the risk of data exfiltration during analysis.
  • VPC Endpoints: Private communication with necessary AWS services (Amazon S3, AWS Systems Manager, AWS KMS) is conducted strictly through AWS PrivateLink interface endpoints and S3 gateway endpoints governed by strict endpoint policies.
  • Forensic Workstations: Pre-configured EC2 instances equipped with specialized forensic toolkits (e.g., memory analysis frameworks, disk inspection tools, packet decoders) mounted to cloned snapshots of compromised production volumes.

Pre-Provisioned Emergency Cross-Account IAM Roles & Break-Glass Procedures

During a major security incident, standard administrative access paths may become inaccessible. For instance, if an enterprise identity provider (IdP) experiences an outage, or if federated Single Sign-On (SSO) credentials have been compromised, incident responders must possess a reliable, pre-authorized mechanism to access spoke accounts immediately.

Break-Glass Role Architecture

Every member account must contain a pre-deployed IAM role specifically designated for emergency incident response, typically named EmergencyIncidentResponderRole. Attempting to create an IAM role during an ongoing incident is too slow, prone to syntax errors, and may be actively blocked if local IAM permissions are corrupted.

Key design requirements for emergency break-glass roles include:

  1. Cross-Account Trust Policy: The trust policy must permit assumption strictly from designated IAM principals in the central Security Tooling Account, rather than allowing arbitrary external principals.
  2. MFA Enforcement: The trust policy must contain a condition checking aws:MultiFactorAuthPresent: "true". An engineer cannot assume the role using plain API keys; they must authenticate with a hardware or virtual MFA device.
  3. Session Duration Limits: Set the role's MaxSessionDuration to a short window (1 hour is the minimum), and require recent MFA with a NumericLessThan condition on aws:MultiFactorAuthAge in the trust policy, as shown below.
  4. Permissions Boundaries: To prevent responders from inadvertently breaking defensive infrastructure, an IAM permissions boundary must be attached that explicitly denies destructive actions against audit logging and security agents.

Example: Emergency Break-Glass Trust Policy

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "EnforceMFAForSecurityAccountAssumption",
      "Effect": "Allow",
      "Principal": {
        "AWS": "arn:aws:iam::111122223333:root"
      },
      "Action": "sts:AssumeRole",
      "Condition": {
        "Bool": {
          "aws:MultiFactorAuthPresent": "true"
        },
        "NumericLessThan": {
          "aws:MultiFactorAuthAge": "3600"
        }
      }
    }
  ]
}

Securing Break-Glass Credentials in AWS Secrets Manager

If a break-glass scenario requires an emergency IAM user (for situations where all role assumption infrastructure is suspect):

  • The user's credentials must never be distributed to individual engineers or stored on local laptops.
  • Secrets must be generated with high-entropy passwords, stored in AWS Secrets Manager in the central Security Tooling account, and locked behind a strict resource-based policy.
  • An Amazon EventBridge rule monitors CloudTrail for the GetSecretValue API call targeting the break-glass secret, immediately triggering a high-priority Amazon SNS alert that pages the entire security leadership team.

Mitigating Blast Radius via Network and Account Boundaries

Blast radius represents the maximum extent of damage an adversary can inflict if a specific identity, application, or network node is compromised. In AWS, blast radius mitigation is enforced through multiple architectural layers:

LayerIsolation MechanismPrimary Blast Radius Mitigation
Organization / AccountService Control Policies (SCPs)Prevents local admins from disabling security tools or altering logs
Network (Global)Route 53 Resolver DNS FirewallBlocks C2 domain resolution across all VPCs enterprise-wide
Network (Inter-VPC)Transit Gateway Route QuarantineBlackholes or isolates compromised VPC traffic from corporate networks
Network (Subnet)Network Access Control Lists (NACLs)Statelessly drops all ingress and egress traffic for an entire subnet
Network (Workload)Quarantine Security GroupsImmediately severs all stateful TCP/UDP connections to an instance
IdentitySession Revocation & Inline DenyInvalidates active STS temporary tokens and blocks lateral API calls

Service Control Policies (SCPs) as Guardrails

Service Control Policies (SCPs) managed in AWS Organizations act as unbreakable filters above individual IAM policies. An explicit deny in an SCP overrides all identity-based policies, resource-based policies, and even actions performed by the root user of a member account. To prevent an adversary from destroying evidence or disabling defense mechanisms, apply an SCP across all workload accounts that explicitly denies tampering with security controls:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyDisablingSecurityServices",
      "Effect": "Deny",
      "Action": [
        "cloudtrail:StopLogging",
        "cloudtrail:DeleteTrail",
        "guardduty:DeleteDetector",
        "guardduty:DisassociateFromMasterAccount",
        "securityhub:DisableSecurityHub",
        "config:DeleteConfigRule"
      ],
      "Resource": "*"
    },
    {
      "Sid": "DenyModifyingForensicRoles",
      "Effect": "Deny",
      "Action": [
        "iam:DeleteRole",
        "iam:DeleteRolePolicy",
        "iam:PutRolePolicy",
        "iam:DetachRolePolicy"
      ],
      "Resource": "arn:aws:iam::*:role/EmergencyIncidentResponderRole"
    }
  ]
}

Network Containment Strategies

When a compute instance (EC2 or container host) is compromised, incident responders must immediately isolate it without destroying volatile memory (RAM) or alerting the attacker through a sudden reboot:

  1. Quarantine Security Group: Security groups are stateful, and tracked connections are not interrupted when rules change; they keep flowing until they time out. AWS guidance is therefore two steps: first move the instance to a security group that allows all traffic in and out (0.0.0.0/0), which makes existing flows untracked; then swap to an isolation security group with no inbound and no outbound rules. Untracked flows stop immediately when the rule that allowed them disappears. The managed runbook AWS-QuarantineEC2Instance assigns such a no-traffic group, and a subnet NACL deny stops traffic instantly regardless of tracking state.
  2. Network ACL (NACL) Rules: Because NACLs are stateless and operate at the subnet boundary, adding rule number 1 with an explicit DENY for all traffic (0.0.0.0/0) drops every packet entering or leaving the subnet, shielding other subnets in the VPC from lateral spread.
  3. Transit Gateway Route Isolation: If an entire VPC is suspected of deep compromise (e.g., active worm propagation), the security team can disassociate the VPC's attachment from the default AWS Transit Gateway (TGW) route table and associate it with a dedicated Quarantine Route Table containing no routes to other workload VPCs, on-premises Direct Connect gateways, or shared internet egress points.

AWS Shield Advanced Proactive Engagement & DRT Delegation

Distributed Denial of Service (DDoS) attacks pose a unique operational challenge because they can overwhelm application availability and exhaust administrative capacity. While AWS Shield Standard provides automatic, baseline Layer 3 and Layer 4 protection at no additional cost (defending against SYN floods, UDP reflection attacks, and network amplification), mission-critical workloads require AWS Shield Advanced.

The AWS Shield Response Team (SRT)

AWS Shield Advanced subscribers with Business or Enterprise Support gain 24/7 access to the AWS Shield Response Team (SRT), formerly called the DDoS Response Team (DRT). During an active volumetric or application-layer attack, customers can engage the SRT to analyze attack signatures and write specialized mitigation rules directly into the customer's AWS WAF web ACLs.

Proactive Engagement Workflow

Under high stress, manually opening an AWS Support ticket and escalating to the SRT introduces unacceptable latency. AWS Shield Advanced provides a Proactive Engagement feature that reverses this workflow: when an attack is detected, the SRT proactively contacts the customer and begins mitigating the attack.

To enable Proactive Engagement, the following components must be configured:

  1. Amazon Route 53 Health Checks: The customer must configure Route 53 health checks and associate them directly with the protected resources (CloudFront distributions, Application Load Balancers, or Elastic IPs). The SRT uses these health checks to verify that an increase in traffic is actually causing service degradation rather than representing a legitimate marketing traffic spike.
  2. Emergency Contact Details: Up to 10 proactive contacts (with phone numbers and email addresses) defined in the Shield Advanced console, including on-call SOC escalation paths.
  3. SRT access role: Let Shield create a role (or choose an existing one) that trusts the drt.shield.amazonaws.com service principal and has the AWS managed policy AWSShieldDRTAccessPolicy. This lets the SRT call Shield Advanced and AWS WAF APIs on your behalf and read your AWS WAF web ACL logs.
  4. Optional extra log buckets: WAF web ACL logs need no extra step. For other evidence (ALB or CloudFront logs, packet captures), add up to 10 buckets and Shield grants the SRT s3:GetBucketLocation, s3:GetObject, and s3:ListBucket. Those buckets must be unencrypted or SSE-S3; the SRT cannot read KMS-encrypted buckets.
  5. Support plan: The account must have Business Support or Enterprise Support.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowDRTAssumeRole",
      "Effect": "Allow",
      "Principal": {
        "Service": "drt.shield.amazonaws.com"
      },
      "Action": "sts:AssumeRole"
    }
  ]
}

Comparison of Incident Containment Strategies

StrategyScopeImplementationSpeedStatefulnessImpact on Volatile Memory
Quarantine Security GroupSingle Instance / ENIAllow-all SG first, then a no-rule isolation SGSecondsStateful (tracked connections persist unless first made untracked)Preserves RAM completely
Network ACL Deny RuleEntire SubnetRule #1 DENY 0.0.0.0/0 ingress/egressMillisecondsStateless (drops all packets at subnet boundary)Preserves RAM completely
Transit Gateway Route IsolationEntire VPCDisassociate from TGW route tableSecondsRouting Layer (drops inter-VPC / hybrid traffic)Preserves RAM completely
Stop EC2 InstanceSingle InstanceAPI: ec2:StopInstances1-2 minutesControl PlaneDestroys volatile RAM completely
Apply Deny SCPAccount or OUOrganization Policy UpdateSecondsIAM / Control Plane (blocks all matching API calls)No direct instance memory impact

Exam Tips & Common Traps

Important

Preserve Volatile Memory: Exam scenarios frequently ask how to contain an EC2 instance that is actively exfiltrating data or mining cryptocurrency while preserving forensic evidence. Never choose to stop, reboot, or terminate the instance. Stopping or rebooting an EC2 instance immediately wipes its volatile RAM, destroying loaded kernel modules, unencrypted cryptographic keys, and active network socket tables. The correct approach is always network quarantine via Security Group modification, followed by live memory acquisition (using SSM Run Command or kernel tools), and then taking EBS volume snapshots.

Warning

Common Trap on Break-Glass Root Access: Questions often present an option recommending logging in as the AWS account root user with a shared password to handle emergency incident response. This is always an incorrect distracter. AWS security best practices mandate that root user credentials must have MFA enabled, root access keys must be deleted, and emergency access must use pre-provisioned cross-account IAM roles with MFA conditions (aws:MultiFactorAuthPresent: "true").

Tip

Shield Proactive Engagement Prerequisites: Remember that Shield Advanced Proactive Engagement will NOT trigger unless Amazon Route 53 health checks are explicitly associated with the protected resource. Without associated health checks, the SRT will not initiate contact because they cannot programmatically differentiate an aggressive DDoS attack from a legitimate viral traffic surge.

Loading diagram...
Multi-Account Incident Readiness & Blast Radius Containment Architecture
Test Your Knowledge

A security engineer detects that an Amazon EC2 instance in a critical production account is actively communicating with a known malicious command-and-control (C2) server. The organization must prevent the instance from transmitting further data while acquiring forensic artifacts for investigation. Which sequence of actions best mitigates the blast radius while preserving evidentiary integrity?

A

Move the instance to an allow-all security group so existing flows become untracked, then to a quarantine security group with no inbound or outbound rules; capture volatile memory with an out-of-band Systems Manager command; snapshot all attached EBS volumes; and share the snapshots with an isolated Forensics account.

B

Immediately execute an ec2:StopInstances API call to freeze disk state, detach the primary elastic network interface (ENI), and launch a replacement instance from the most recent daily backup snapshot.

C

Reboot the instance into single-user rescue mode, install open-source forensic memory imaging software, upload the memory dump to an S3 bucket in the production account, and terminate the instance.

D

Apply an organization-wide Service Control Policy (SCP) denying ec2:* actions across the production account, detach all EBS volumes from the running instance, and export them directly to the Security Tooling account.

Test Your Knowledge

An enterprise subscribes to AWS Shield Advanced (with Business Support) to protect its mission-critical web applications hosted behind Application Load Balancers (ALBs). The Chief Information Security Officer (CISO) wants the AWS Shield Response Team (SRT) to proactively contact on-call engineers and help implement mitigation rules in AWS WAF during Layer 7 DDoS attacks, without requiring the internal team to manually open a support case. Which combination of configurations is mandatory to fulfill this requirement?

A

Enable AWS Shield Standard, deploy AWS Network Firewall in front of the ALBs, and configure an Amazon SNS topic subscribed to the AWS Support email endpoint.

B

Configure an Amazon CloudWatch anomaly detection alarm on ALB request count, grant AWS Support administrator access to the AWS Organizations root, and enable WAF rate-limiting.

C

Associate Amazon Route 53 health checks with the protected ALBs, enable proactive engagement with at least one contact, and grant SRT access through a role that trusts drt.shield.amazonaws.com with the AWSShieldDRTAccessPolicy managed policy.

D

Deploy an Amazon API Gateway in front of the ALBs, create a custom Lambda authorizer that pushes blocked IP addresses to the SRT, and configure AWS Config to monitor ALB security groups.

Test Your Knowledge

A security architect is establishing an emergency break-glass procedure for responders during severe incidents where corporate federated Single Sign-On (SSO) is completely unavailable. Which architecture best adheres to AWS security best practices for emergency incident response?

A

Store the AWS account root user credentials in an unencrypted shared team spreadsheet accessible by all operations staff, and disable root MFA to avoid delay during emergencies.

B

Pre-provision an emergency cross-account IAM role in every spoke account that trusts the centralized Security Tooling account, enforces multi-factor authentication using the aws:MultiFactorAuthPresent condition key, restricts session duration to one hour, and alerts the SOC via EventBridge whenever assumed.

C

Create a local IAM user named 'break-glass-admin' in each spoke account with AdministratorAccess and active access keys, committing the secret access keys to a password-protected internal code repository.

D

Deploy an AWS Systems Manager hybrid activation code across all production EC2 instances so engineers can open remote interactive shell sessions without requiring any IAM permissions or network connectivity.

Test Your Knowledge

A multinational enterprise uses AWS Organizations with hundreds of member accounts. During an internal security audit, the security team discovers that a local administrator in a development account disabled AWS CloudTrail to conceal unauthorized resource provisioning. Which governance mechanism guarantees that administrators in member accounts cannot disable CloudTrail, Amazon GuardDuty, or AWS Security Hub under any circumstances?

A

Configure IAM permissions boundaries on every IAM user and role created within the development account to explicitly deny CloudTrail modification APIs.

B

Attach an AWS WAF rule group to the AWS Management Console login URL that inspects request bodies and blocks any calls containing security deactivation actions.

C

Deploy an Amazon EventBridge rule in the development account that triggers an AWS Lambda function to re-enable CloudTrail within 60 seconds of being stopped.

D

Apply a Service Control Policy (SCP) at the Root or Organizational Unit (OU) level that explicitly denies cloudtrail:StopLogging, cloudtrail:DeleteTrail, guardduty:DeleteDetector, and securityhub:DisableSecurityHub.

Sections you finish are checked off in the contents.