3.2 Testing & Validating Incident Response Plans

Key Takeaways

  • AWS Fault Injection Service (FIS) executes controlled chaos engineering experiments against production-like environments to validate security detection, alerting thresholds, and automated response runbooks.

  • FIS experiment templates define specific target resources, disruption actions (such as CPU stress, network latency, and process termination), and automated CloudWatch alarm stop conditions that immediately abort the experiment if safety thresholds are breached.

  • AWS Resilience Hub assesses application components against defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets, uncovering architectural drift and generating tailored FIS experiment templates.

  • Tabletop exercises validate decision-making hierarchies, legal disclosure requirements, and cross-functional communications, whereas technical GameDays validate operational detection telemetry and automated containment tools under live conditions.

  • Runbook drift occurs when infrastructure modifications, IAM policy updates, or API changes render operational playbooks obsolete; continuous automated testing is required to identify and remediate these procedural gaps.

Last updated: September 2026

Testing & Validating Incident Response Plans

Having documented incident response playbooks is insufficient for modern cloud security. During an actual security breach, organizations frequently discover that runbooks are out of date, automation scripts fail due to missing IAM permissions, contact lists contain former employees, or security alerting thresholds fail to fire under real-world conditions. Incident readiness must be validated empirically. Through controlled fault injection, automated resiliency assessments, and multi-team simulation exercises, security teams uncover operational blind spots and eliminate runbook drift before adversaries strike.


The Imperative for Incident Response Testing: Beyond Paper Runbooks

Traditional disaster recovery often relied on periodic, scheduled documentation reviews. In dynamic cloud environments, infrastructure changes continuously through infrastructure as code (IaC) pipelines, microservice updates, and automated scaling. This rapid pace introduces runbook drift—a condition where documented standard operating procedures (SOPs) no longer reflect the technical reality of the running infrastructure.

To counter runbook drift, organizations adopt Security Chaos Engineering:

+-------------------------------------------------------------+
| 1. Formulate Hypothesis:                                    |
|    "Terminating the security agent will trigger an alert   |
|     and initiate auto-remediation within 2 minutes."        |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| 2. Inject Controlled Disruption (AWS FIS):                  |
|    Kill the security monitoring daemon on 20% of instances.|
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| 3. Monitor Telemetry & Safety Guardrails:                   |
|    Verify CloudWatch metrics; stop if 5xx error rate > 5%.  |
+-------------------------------------------------------------+
                               |
                               v
+-------------------------------------------------------------+
| 4. Validate Outcome & Remediate Gaps:                       |
|    Did the alert fire? Did the runbook execute successfully?|
+-------------------------------------------------------------+

By systematically testing hypotheses about how defensive systems respond to failure or attack, teams build empirical confidence in their detection and response capabilities.


AWS Fault Injection Service (FIS) Mechanics & Security Testing

AWS Fault Injection Service (FIS) is a fully managed chaos engineering service designed to run controlled fault injection experiments on AWS workloads. While frequently used to test high availability and fault tolerance, FIS is an essential tool for validating incident response and security monitoring pipelines.

Core Components of AWS FIS

An FIS experiment is defined by a declarative Experiment Template consisting of four primary building blocks:

  1. Targets: The specific AWS resources upon which the experiment executes. Targets can be Amazon EC2 instances, Amazon ECS tasks, Amazon EKS pods, Amazon RDS databases, or IAM roles. Targets are resolved dynamically using resource tags, resource filters, or specific ARNs. Responders define target scopes using selectionMode: "PERCENT(n)" (e.g., impact 20% of matching instances) or selectionMode: "COUNT(n)" (e.g., impact exactly 1 instance) to strictly control blast radius.
  2. Actions: The precise fault or disruption injected into the targets. Actions can execute sequentially or concurrently. AWS FIS provides pre-packaged action plugins for compute disruption, network impairment, and Systems Manager command execution.
  3. Stop Conditions: Optional in the template syntax (a template can specify none) but essential for any production-like experiment. Stop conditions are linked directly to Amazon CloudWatch alarms. If an experiment causes unintended collateral damage—such as customer-facing 5xx error rates spiking, synthetic transaction latency exceeding thresholds, or unhandled exceptions—the linked CloudWatch alarm transitions to the ALARM state. FIS stops the experiment; a stopped experiment cannot be resumed.
  4. Service Role: An IAM service role assumed by fis.amazonaws.com that grants FIS the explicit permissions required to manipulate the target resources and invoke Systems Manager documents.

Specific FIS Experiments for Security Readiness

Security teams configure specialized FIS actions to validate detection and response runbooks:

1. Terminating Security and Endpoint Protection Agents

Adversaries often attempt to disable host-based security agents (such as antivirus daemons, file integrity monitors, or the Amazon CloudWatch agent) upon compromising an instance. Responders use the aws:ssm:send-command FIS action to target running instances and execute the AWSFIS-Run-Kill-Process SSM document:

{
  "description": "Simulate Malicious Termination of Endpoint Security Agent",
  "targets": {
    "AppInstances": {
      "resourceType": "aws:ec2:instance",
      "resourceTags": {
        "Environment": "Staging",
        "Tier": "Application"
      },
      "selectionMode": "COUNT(1)"
    }
  },
  "actions": {
    "KillSecurityAgent": {
      "actionId": "aws:ssm:send-command",
      "parameters": {
        "documentArn": "arn:aws:ssm:us-east-1::document/AWSFIS-Run-Kill-Process",
        "documentParameters": "{\"ProcessName\": \"security-agent-daemon\"}",
        "duration": "PT5M"
      },
      "targets": {
        "Instances": "AppInstances"
      }
    }
  },
  "stopConditions": [
    {
      "source": "aws:cloudwatch:alarm",
      "value": "arn:aws:cloudwatch:us-east-1:123456789012:alarm:Application5xxErrorRateExceeded"
    }
  ],
  "roleArn": "arn:aws:iam::123456789012:role/FISSecurityExperimentRole"
}

Validation Goal: Does the absence of heartbeats trigger a CloudWatch alarm? Does an automated Systems Manager runbook restart the agent or quarantine the instance?

2. Network Disruption & Isolation Simulation

Using the aws:network:disrupt-connectivity action, FIS blocks traffic for target subnets (all traffic, or traffic to a specific destination such as Amazon S3 or DynamoDB). Latency and packet-loss faults come from SSM-based actions such as AWSFIS-Run-Network-Latency. This validates whether out-of-band security logging continues streaming when an application's primary egress path is degraded.

3. Resource Starvation (CPU & Memory Spikes)

Adversaries running cryptomining payloads saturate CPU cores. Using AWSFIS-Run-CPU-Stress, FIS forces 100% CPU utilization across target instances. This verifies that CPU alarms, anomaly detection bands, and Auto Scaling or containment workflows respond correctly under heavy load. It does not produce GuardDuty cryptocurrency findings, which depend on network and DNS indicators such as mining-pool lookups.

4. Disabling IAM Credentials and Simulating Policy Revocation

Security runbooks often involve revoking IAM credentials or modifying security group rules during containment. FIS experiments can simulate permission loss or security group changes to verify whether applications degrade gracefully and whether alerting systems notify the SOC of administrative tampering.


AWS Resilience Hub: Assessing Resiliency Policies & Drift

While FIS actively injects faults, AWS Resilience Hub provides a centralized control plane to define, assess, and track the resilience of cloud applications against declared business targets.

+-----------------------+      +-----------------------+      +-----------------------+
| 1. Model Application  | ---> | 2. Define Policy      | ---> | 3. Assess & Discover  |
| • CloudFormation      |      | • Target RTO          |      | • Calculate Est. RTO  |
| • AppRegistry         |      | • Target RPO          |      | • Identify Gaps/Drift |
+-----------------------+      +-----------------------+      +-----------------------+
                                                                          |
                                                                          v
+-----------------------+      +-----------------------+      +-----------------------+
| 6. Re-assess & Track  | <--- | 5. Run FIS Chaos Test | <--- | 4. Generate Artifacts |
| • Continuous Baseline |      | • Validate Hypotheses |      | • FIS Templates       |
+-----------------------+      +-----------------------+      | • Standard Ops (SOPs) |
+-----------------------+

1. Application Modeling

Resilience Hub discovers and groups application components from multiple AWS sources, including AWS CloudFormation stacks, Terraform state files, AWS Service Catalog AppRegistry, or resource tags. It decomposes the application into functional components (compute, database, storage, networking).

2. Defining Resiliency Policies

Organizations establish explicit business targets for two foundational metrics:

  • Recovery Time Objective (RTO): The maximum acceptable duration of downtime following a disruption before service must be restored.
  • Recovery Point Objective (RPO): The maximum acceptable data loss measured in time (e.g., no more than 5 minutes of transactional data lost).

Resilience Hub evaluates these metrics across four distinct disruption scopes:

  1. Application Component: Software bugs, container crashes, or unhandled exceptions.
  2. Infrastructure: Virtual machine hardware failure, underlying EBS volume degradation.
  3. Availability Zone (AZ): Datacenter failure or total AZ connectivity loss.
  4. Region: Comprehensive disaster recovery across geographic areas.

3. Assessment & Gap Identification

When an assessment executes, Resilience Hub compares the actual architectural configuration against the declared resiliency policy. It calculates an Estimated RTO and Estimated RPO.

  • If a critical relational database is configured as a Single-AZ instance, but the AZ policy requires an RTO of 15 minutes, Resilience Hub flags a policy breach because failing over a Single-AZ database manually requires hours.
  • It detects architectural drift: when developers deploy new unbacked-up S3 buckets or unclustered compute instances, Resilience Hub surfaces these components as unmanaged risks.

4. Generating FIS Experiment Recommendations

Crucially, Resilience Hub does not just generate static reports; it exports actionable operational code:

  • Standard Operating Procedures (SOPs): Pre-built Systems Manager Automation runbooks to execute failover or recovery procedures.
  • Recommended FIS Experiment Templates: Resilience Hub generates ready-to-deploy AWS FIS experiment templates tailored specifically to test the failure modes evaluated during the assessment. Running these experiments empirically proves whether the system satisfies its declared RTO/RPO targets.

GameDays and Tabletop Exercises

Validating incident readiness requires exercising both technical automation and human decision-making. Organizations employ two complementary simulation methodologies:

Tabletop Exercises (Discussion-Based)

A Tabletop Exercise (TTX) is a structured, scenario-driven discussion bringing together cross-functional stakeholders—including security engineers, legal counsel, compliance officers, human resources, public relations, and executive leadership—in a collaborative setting.

  • Scope: Evaluates decision-making hierarchies, legal disclosure requirements, communication protocols, and executive escalations.
  • Scenario Examples: A major ransomware attack threatening public data exfiltration; an insider threat leaking AWS root account credentials; discovery of an active zero-day vulnerability in a third-party dependency requiring emergency patching.
  • Key Questions Explored:
    • Who has the legal authority to shut down a revenue-generating production system to contain an attack?
    • What are the regulatory reporting deadlines (e.g., GDPR 72-hour notification, SEC Form 8-K 4-day disclosure rule for material cyber incidents)?
    • How do we communicate with customers and law enforcement without tipping off the adversary?

Technical GameDays (Live-Fire Simulations)

A Technical GameDay is a hands-on, live-action operational exercise conducted against a production-like staging environment (or controlled production slice). During a GameDay, engineers face realistic, unannounced technical failures or synthetic adversary attacks.

FeatureTabletop Exercise (TTX)Technical GameDayAutomated FIS Experiment
Primary ObjectiveValidate governance, communication, and decision policyValidate technical runbooks, detection telemetry, and containmentValidate infrastructure self-healing and alert thresholds
ParticipantsCross-functional leaders (Legal, PR, Execs, SecOps)Hands-on responders (SOC, DevOps, Cloud Engineers)Fully automated (FIS service and CloudWatch)
EnvironmentConference room / virtual discussion (No live infrastructure)Staging or isolated production sandbox environmentStaging or production environments with stop conditions
Duration2 to 4 hoursHalf-day to full-day eventMinutes (bounded by experiment duration)
Artifact ProducedUpdated escalation policies, legal playbooks, communication treesVerified runbooks, patched IAM policies, updated alert rulesAutomated pass/fail metrics, RTO/RPO validation data

The Red, Blue, and White Team Model

Technical GameDays typically organize participants into three structured teams:

  • Red Team (Adversary Simulation): Injects synthetic attacks, runs FIS templates, terminates security processes, or simulates credential exfiltration. The Red Team must adhere to strict rules of engagement.
  • Blue Team (Defenders): The active incident response engineers and SOC analysts who must detect, investigate, and contain the simulated intrusion using existing tools (CloudTrail, GuardDuty, Security Hub, Athena, runbooks).
  • White Team (Referees & Observers): Impartial facilitators who define the exercise boundaries, monitor safety stop conditions, inject situational clues if defenders become completely stuck, and record objective timelines for the debrief.

Identifying Runbook Drift & Operational Gaps

When a GameDay or FIS experiment concludes, the most valuable output is the catalog of discovered operational failures. Common operational gaps surfaced during testing include:

  1. IAM Permission Decay: An automated remediation runbook fails because an administrator updated an IAM permissions boundary or SCP three months earlier, inadvertently revoking the runbook's ability to call ec2:ModifyInstanceAttribute or s3:PutBucketPublicAccessBlock.
  2. Stale Notification Routing: Alert notifications fail to deliver because an Amazon SNS topic is subscribed to an unmonitored distribution list or SMS phone numbers belonging to departed personnel.
  3. EventBridge Schema Mismatches: An automated remediation Lambda function fails because an upstream AWS service updated its finding JSON format, breaking a hardcoded JSON path parser in the custom script.
  4. Telemetry Ingestion Lag: The Blue Team takes 45 minutes to detect an attack because CloudTrail management event delivery to S3 experienced normal propagation delay, revealing that the team should have configured real-time CloudWatch Logs metric filters or EventBridge rules instead of relying on batch S3 queries.

Operationalizing Continuous Validation

To ensure runbooks do not drift, forward-leaning organizations embed chaos testing directly into their CI/CD deployment pipelines:

  • When a cloud engineering team updates a Terraform or CloudFormation stack, an automated staging pipeline deploys the infrastructure, triggers an AWS FIS experiment to simulate security agent failure, verifies that the SOC alerting pipeline receives the event within 120 seconds, and aborts the release if validation fails.

Exam Tips & Common Traps

Important

FIS Stop Conditions Are the Safety Net: Exam scenarios frequently ask how to safely execute chaos testing in production without risking widespread outages. The key mechanism is the CloudWatch alarm stop condition. If an option mentions running FIS experiments without stop conditions or relying on manual console monitoring to stop an experiment, it is incorrect. FIS stops the experiment as soon as a stop-condition alarm enters the ALARM state.

Warning

Common Trap on Resilience Hub Capabilities: Do not confuse AWS Resilience Hub with AWS Backup or AWS Fault Injection Service. Resilience Hub does not perform data backups, and it does not directly execute chaos attacks. Instead, Resilience Hub assesses application architectures against declared RTO/RPO policies, detects drift, and generates recommended FIS experiment templates and SSM automation SOPs.

Tip

Selecting the Right Exercise Type: On the SCS-C03 exam, pay close attention to what the question asks you to validate:

  • If the scenario requires validating executive communication, regulatory breach notification timelines (e.g., SEC/GDPR), and legal authorities, choose a Tabletop Exercise.
  • If the scenario requires validating real-time alerting, engineer triage speed, and hands-on runbook execution under simulated adversary pressure, choose a Technical GameDay.
  • If the scenario requires continuous, programmatic validation of infrastructure resilience and automated failover bounded by strict error rate alarms, choose AWS Fault Injection Service (FIS).
Loading diagram...
AWS FIS Experiment Lifecycle & Resilience Hub Feedback Loop
Test Your Knowledge

A cloud security team wants to verify that their monitoring pipeline detects when an endpoint detection and response (EDR) daemon is terminated on production EC2 instances. They configure an AWS Fault Injection Service (FIS) experiment targeting a staging environment. Which mechanism guarantees that the experiment will halt automatically if the application's overall HTTP 5xx error rate exceeds 5% during testing?

A

Attach an AWS WAF rate-based rule to the Application Load Balancer that blocks traffic from the FIS service role when error rates exceed 5%.

B

Configure an Amazon CloudWatch metric alarm that monitors HTTP 5xx error percentages and define that alarm as a stop condition within the FIS experiment template.

C

Create a Service Control Policy (SCP) at the organizational unit level that revokes the FIS service role's permissions when CloudWatch Logs detects an error threshold breach.

D

Embed an AWS Lambda script inside the FIS experiment action definition that continuously polls application logs and issues an abort API call.

Test Your Knowledge

An enterprise security architect must validate whether a multi-tier banking application can achieve a 15-minute Recovery Time Objective (RTO) during an Availability Zone outage. Which AWS service should the architect use to assess the application architecture against this target RTO policy, identify single-point-of-failure drift, and generate tailored chaos engineering experiment templates?

A

AWS Systems Manager OpsCenter

B

AWS Trusted Advisor

C

AWS Security Hub

D

AWS Resilience Hub

Test Your Knowledge

A healthcare provider is preparing its annual incident response readiness evaluation. The Chief Information Security Officer (CISO) wants to assess executive escalation procedures, regulatory breach disclosure requirements under HIPAA and state privacy laws, and public relations crisis messaging, without injecting any technical failures into cloud workloads. Which exercise format is most appropriate for this objective?

A

A structured Tabletop Exercise (TTX) involving legal counsel, executive leadership, communications directors, and security managers walking through a simulated patient record breach scenario.

B

An automated AWS Fault Injection Service (FIS) chaos experiment executed against the production database cluster during business hours.

C

An unannounced technical Red Team penetration test attempting to extract production database snapshots to an external public IP address.

D

A hands-on Technical GameDay where junior cloud engineers attempt to recover an encrypted database from cold backup storage.

Test Your Knowledge

During a simulated incident response exercise, an automated Systems Manager runbook fails to isolate a compromised EC2 instance because an IAM policy update applied three months earlier removed the runbook role's permission to call ec2:ModifyInstanceAttribute. This failure is a prime example of which operational risk, and how should it be systematically prevented?

A

Compute starvation risk; prevented by configuring EC2 Auto Scaling groups with dynamic target tracking scaling policies.

B

Log ingestion latency; prevented by configuring Amazon Security Lake to ingest VPC Flow Logs using the Open Cybersecurity Schema Framework (OCSF).

C

Runbook drift; prevented by establishing periodic automated testing of incident response playbooks and integrating FIS chaos experiments into continuous deployment pipelines.

D

Blast radius leakage; prevented by applying an organization-wide Service Control Policy (SCP) that denies all IAM policy modifications across every account.

Sections you finish are checked off in the contents.