1.1 Workload Monitoring & Health Check Strategies
Key Takeaways
CloudWatch metric alarms evaluate threshold breaches over defined periods, while composite alarms evaluate Boolean expressions across multiple alarms to eliminate alert fatigue and detect compound attack patterns.
CloudWatch Metric Math enables dynamic baselining and statistical deviation alerts using functions like RATE, DIFF, and ANOMALY_DETECTION_BAND without hardcoding static thresholds.
CloudWatch Synthetics canaries run automated Node.js or Python scripts inside AWS Lambda to continuously validate endpoint integrity, TLS certificates, HTTP status codes, and DOM elements from outside or inside private VPCs.
Route 53 health checks support endpoint probing, calculated combinators, and CloudWatch alarm monitoring to drive automated DNS failover and quarantine compromised application targets.
Amazon EventBridge rule patterns match structured JSON event payloads from CloudWatch, GuardDuty, and AWS API calls, routing actionable security telemetry to Lambda, SQS, or Step Functions with dead-letter queue (DLQ) resiliency.
1.1 Workload Monitoring & Health Check Strategies
Security monitoring in AWS requires moving beyond traditional operational health metrics to establish continuous behavioral visibility across compute, networking, and application layers. Security-focused workload monitoring detects unauthorized access, resource exploitation, denial-of-service attempts, and data exfiltration patterns in near real time. A comprehensive detection architecture combines Amazon CloudWatch metrics and composite alarms, CloudWatch Metric Math, CloudWatch Synthetics canaries, Amazon Route 53 health checks, and Amazon EventBridge event-driven routing.
CloudWatch Metrics, Metric Math & Anomaly Detection
CloudWatch serves as the foundational metrics repository in AWS. While operational engineering teams track CPU utilization and disk space, security engineers analyze metrics to detect indicators of compromise (IoC) and abnormal workload behavior.
Metric Evaluation Windows and High-Resolution Metrics
Standard CloudWatch metrics publish at a 1-minute or 5-minute frequency. For mission-critical security boundaries, high-resolution custom metrics provide 1-second, 5-second, 10-second, or 30-second granularity. High-resolution metrics allow rapid detection of short-lived brute-force spikes, token exhaustion attacks, or rapid API throttling events.
When configuring metric alarms, three parameters govern evaluation:
- Period: The length of time associated with each individual datapoint (for example, 60 seconds).
- Evaluation Periods (N): How many of the most recent periods CloudWatch looks at.
- Datapoints to Alarm (M): How many of those N periods must breach. The M datapoints do not need to be consecutive, so alarming when 3 out of the last 5 periods breach filters out transient network jitter while still catching persistent anomalous activity.
Security Anomaly Detection via Metric Math
Static thresholds often fail in security contexts due to diurnal workload cycles, batch processing schedules, and seasonal traffic swings. A hardcoded threshold of 1,000 HTTP 403 errors per minute may trigger false alerts during peak business hours while missing a credential-stuffing attack occurring during the middle of the night.
CloudWatch Metric Math allows security engineers to execute statistical queries and transformations across multiple metrics:
// Expression to calculate the ratio of HTTP 4xx client errors against total request volume
// m1 = HTTPCode_Target_4XX_Count (Application Load Balancer)
// m2 = RequestCount (Application Load Balancer)
e1 = (m1 / m2) * 100
Beyond basic arithmetic, CloudWatch offers built-in Anomaly Detection. Anomaly detection applies machine learning algorithms to continuous metric history, automatically establishing dynamic upper and lower prediction bands that account for hourly, daily, and weekly seasonality. Security engineers configure alarms to fire when metrics breach these dynamic confidence bands:
// Detect when outbound network egress exceeds expected upper confidence band by 3 standard deviations
ANOMALY_DETECTION_BAND(m1, 3)
Missing Data Treatment in Security Alarms
A critical exam consideration is how an alarm handles missing datapoints. CloudWatch provides four options:
| Setting | Behavior | Security Operational Context |
|---|---|---|
missing | Alarm transitions to INSUFFICIENT_DATA | Standard default; risky for security if an adversary stops a logging agent. |
ignore | Current state is maintained | Retains previous state; suppresses flapping during known reporting outages. |
breaching | Missing datapoints count as exceeding threshold | Critical for security heartbeats. If an endpoint or security agent stops reporting, the alarm transitions to ALARM immediately. |
notBreaching | Missing datapoints are treated as healthy | Ideal for error-count metrics (e.g., WAF blocked requests) where zero errors results in no emitted datapoints. |
Exam Tip: For heartbeat metrics where a security daemon or canary must report continuously, always configure missing data as
breaching. If an adversary terminates the monitoring process or disables the network interface, the absence of data immediately triggers incident response.
Composite Alarms & Alarm Storm Suppression
In complex architectures, a single incident can generate dozens of correlated alerts. For instance, a distributed denial-of-service (DDoS) attack against an application tier can simultaneously trigger alarms for ALB response latency, target 5xx errors, backend EC2 CPU utilization, and database connection queue depth. Flooding security operations center (SOC) analysts with fragmented notifications induces alert fatigue and delays triage.
CloudWatch Composite Alarms aggregate multiple metric alarms using Boolean logic (AND, OR, NOT). Composite alarms do not evaluate raw metrics directly; they evaluate the state transitions of underlying metric alarms.
Boolean Expression Syntax
A composite alarm rule evaluates child alarm ARNs or alarm names:
ALARM("High-ALB-HTTP-4XX-Ratio")
AND ALARM("High-WAF-Blocked-Requests")
AND NOT ALARM("Scheduled-Penetration-Testing-Active")
This expression ensures that a high-priority SOC alert fires only when both the load balancer and the web application firewall detect malicious patterns concurrently, while suppressing alerts if an operational maintenance or penetration testing alarm is currently active.
Alarm Actions and Suppressor Alarms
Composite alarms can also define an Alarm Suppressor. An alarm suppressor acts as a master gate: while the suppressor alarm is in the ALARM state, the composite alarm stops executing its configured notifications (such as Amazon SNS topics or Systems Manager OpsCenter items). This prevents secondary notification storms when a known major failure (such as an entire Direct Connect link failure) is already under active remediation.
CloudWatch Synthetics Canaries for Endpoint Integrity
While CloudWatch metrics provide internal resource telemetry, security teams require outside-in verification to detect silent endpoint degradation, unauthorized DNS redirection, SSL/TLS certificate tampering, and web defacement.
CloudWatch Synthetics deploys configurable scripts called canaries. Canaries run on a scheduled interval (e.g., every 1 or 5 minutes) powered by AWS Lambda, using Node.js runtimes (Puppeteer or Playwright driving headless Chromium) or Python runtimes (Selenium).
Canary Security Use Cases
- TLS/SSL Handshake & Expiration Auditing: Canaries validate that HTTPS endpoints present valid certificates signed by authorized authorities, catching certificate expirations before public exposure and flagging unauthorized cipher downgrades.
- Security Header Enforcement: Canaries assert the presence of defensive HTTP response headers, including
Strict-Transport-Security(HSTS),Content-Security-Policy(CSP),X-Content-Type-Options: nosniff, andX-Frame-Options: DENY. - DOM Defacement & Malicious Script Injection: Canaries execute automated user journeys, clicking elements and validating that specific DOM nodes or text strings match expected hashes. If an attacker injects unauthorized JavaScript (such as a Magecart digital skimming script), the canary detects altered DOM structures or unexpected third-party outbound network calls.
- Private VPC Canary Execution: For zero-trust internal microservices, canaries can be configured with an Amazon VPC configuration (Subnets and Security Groups). The canary provisions Elastic Network Interfaces (ENIs) inside the private subnet to probe internal ALBs and API Gateways without traversing the public internet. When deployed inside a VPC, the subnet must have access to Amazon S3 (via a VPC Gateway Endpoint) and CloudWatch (via VPC Interface Endpoints or a NAT Gateway) to publish logs, metrics, and screenshots.
Route 53 Health Checks & Automated Failure Isolation
Amazon Route 53 health checks monitor the operational status of web servers, load balancers, and network endpoints from geographically dispersed health checkers around the globe. In security architectures, Route 53 health checks serve as the automated tripwire for quarantine and DNS failover.
Health Check Types
Route 53 provides three distinct health check categories:
- Endpoint Checks: Directly probe an IP address or fully qualified domain name (FQDN) over HTTP, HTTPS, or TCP. Evaluated on either a standard interval (30 seconds) or fast interval (10 seconds), with a configurable failure threshold (typically 3 consecutive failures).
- Calculated Health Checks: Combine up to 256 individual health checks using Boolean logic (
AND,OR,NOT) or an integer threshold rule (e.g., "healthy if at least 3 out of 5 child checks are healthy"). Security architects use calculated checks to monitor multi-tier stacks (e.g., Web tier AND Database tier must both report healthy). - CloudWatch Metric Monitors: Monitor a specific CloudWatch alarm rather than polling an external network endpoint. This allows Route 53 to trigger DNS rerouting based on complex internal application telemetry (e.g., memory exhaustion, authorization failure surges, or WAF block rates) that cannot be probed via external HTTP GET requests.
String Matching for Integrity Verification
A critical feature for security is String Matching on HTTP/HTTPS endpoint checks. Standard HTTP checks merely confirm that the endpoint returns a 2xx or 3xx status code. However, a compromised web server or one experiencing database connection failures might still serve a generic HTTP 200 error page or defaced content.
When string matching is enabled, Route 53 searches the first 5,120 bytes of the HTTP response body for a specific, expected alphanumeric string (e.g., "SYSTEM_HEALTHY_TOKEN_VALID"). If the string is missing—even if the HTTP response code is 200 OK—Route 53 marks the target unhealthy.
DNS Failover Configurations
Route 53 associates health checks with DNS routing policies to automate traffic rerouting:
- Active-Passive Failover: Directs production traffic to a primary resource (e.g., an Application Load Balancer). If the primary health check fails, Route 53 automatically updates DNS resolution to direct queries to a secondary disaster recovery site, a static S3 website containing a read-only emergency incident page, or a network sinkhole.
- Active-Active Routing: Distributes traffic across multiple healthy endpoints using weighted, latency-based, or geolocation routing. Unhealthy endpoints are stripped from DNS responses automatically within seconds.
Exam Tip: Route 53 HTTPS health checks never validate the SSL/TLS certificate, so a check does not fail when the certificate is invalid or expired. Enabling SNI only sends the host name in the TLS
client_helloso the endpoint can present the right certificate; it does not turn on certificate validation. To catch expired or swapped certificates, use a CloudWatch Synthetics canary, ACM expiration events for ACM-managed certificates, or an AWS Config rule.
EventBridge Rule Patterns for Security Event Routing
Amazon EventBridge is the central event bus for security automation in AWS. Rather than polling logs, EventBridge receives events pushed natively from AWS services, custom applications, and SaaS partners in real time.
Event Pattern Grammar & Matching Rules
EventBridge rules evaluate JSON event structures against defined Event Patterns. An event pattern matches only if the fields specified in the pattern are present and match the values in the incoming event.
Key pattern matching capabilities for security engineers include:
- Exact String Matching:
"detail-type": ["AWS API Call via CloudTrail"] - Prefix Matching:
"eventName": [{"prefix": "Delete"}, {"prefix": "Stop"}] - Numeric Value Matching: Matching numerical thresholds, such as GuardDuty severities:
"severity": [{"numeric": [">=", 7]}] - Anything-But Matching: Excluding expected service principals or authorized automated pipelines:
"userIdentity": {"arn": [{"anything-but": "arn:aws:iam::123456789012:role/DeployPipelineRole"}]} - Nested-Field Matching: Flagging sessions that were not MFA-authenticated:
"userIdentity": {"sessionContext": {"attributes": {"mfaAuthenticated": ["false"]}}}
Example Security Event Pattern
The following event pattern triggers an alert whenever an IAM policy is modified or deleted by any entity outside the approved deployment automation role:
{
"source": ["aws.iam"],
"detail-type": ["AWS API Call via CloudTrail"],
"detail": {
"eventSource": ["iam.amazonaws.com"],
"eventName": [
"DeleteAccountPasswordPolicy",
"PutAccountPasswordPolicy",
"AttachGroupPolicy",
"DetachGroupPolicy",
"AttachRolePolicy",
"DetachRolePolicy",
"CreatePolicyVersion",
"SetDefaultPolicyVersion"
],
"userIdentity": {
"sessionContext": {
"sessionIssuer": {
"userName": [{
"anything-but": "CI-CD-Automation-Executor"
}]
}
}
}
}
}
Target Resiliency and Dead-Letter Queues (DLQs)
In security incident response, missed notifications can result in undetected intrusions. When EventBridge routes an event to a target (such as an AWS Lambda remediation function or Amazon SNS topic), the target may fail to process the event due to service throttling, network partitions, or misconfigured IAM permissions.
To ensure auditability and zero event loss:
- Configure Retry Policies: Specify maximum event age (up to 24 hours) and retry attempts (up to 185 attempts) with exponential backoff.
- Configure a Dead-Letter Queue (DLQ) using an Amazon SQS standard queue. Any event that fails delivery after exhausting retries is deposited into the DLQ, where an automated alarm alerts security engineers to investigate the delivery failure.
Specialty Exam Pitfalls & Architectural Traps
- Treating CloudWatch Alarms as Direct Evaluators of Multiple Metrics: Standard metric alarms evaluate a single metric or a single Metric Math output expression. They cannot natively evaluate
MetricA > 10 AND MetricB > 20. To achieve multi-condition evaluations, you must create two individual metric alarms and combine them in a Composite Alarm using Boolean logic. - VPC Endpoint Misconfigurations for Synthetics Canaries: When a Synthetics canary runs inside a private VPC subnet, it cannot access public AWS service endpoints. If the subnet lacks a NAT Gateway or VPC Interface Endpoints for CloudWatch (
com.amazonaws.<region>.monitoring) and S3 Gateway Endpoints (com.amazonaws.<region>.s3), the canary will fail with connection timeout errors and will be unable to upload test artifacts, HAR logs, or screenshots. - Route 53 External Health Check IP Filtering: Route 53 health checkers reside on public AWS IP ranges distributed worldwide. If your network access control lists (NACLs) or EC2 security groups restrict inbound traffic to specific corporate CIDRs without allowing Route 53 health checker IP addresses, health checks will fail even when the workload is healthy.
- EventBridge Default Bus vs Custom Event Buses: Native AWS service events (CloudTrail, GuardDuty, Macie, Health) are emitted only to the default event bus. You cannot route native AWS service events directly onto a custom event bus; you must capture them on the default bus and optionally use a rule to forward them to a custom or cross-account event bus.
A security operations team wants to alert engineers only when an application experiences an active distributed denial-of-service attack. The team wants to trigger an alert if the Application Load Balancer target 5xx error rate exceeds 15% AND the CloudWatch WAF blocked request count exceeds 5,000 per minute, but suppress alerts if a scheduled load testing alarm is currently in ALARM state. Which architectural design satisfies this requirement with the least operational overhead?
Create a single CloudWatch metric alarm using Metric Math that computes the ratio of 5xx errors to total requests, and write an AWS Lambda function to check WAF metrics and load test flags.
Deploy an Amazon Managed Grafana dashboard with custom alerting rules polling CloudWatch APIs every 30 seconds to evaluate the compound conditions.
Configure two underlying CloudWatch metric alarms for the ALB 5xx rate and WAF blocked requests, create a third alarm for load testing, and configure a CloudWatch Composite Alarm evaluating the Boolean expression ALARM(ALB) AND ALARM(WAF) AND NOT ALARM(LoadTesting).
Create an Amazon EventBridge rule that intercepts raw metric streams from CloudWatch, evaluates an inline Python script to calculate the conjunction of conditions, and publishes to Amazon SNS.
A financial services firm runs an internal microservice hosted in private subnets behind an internal Application Load Balancer. The security team mandates continuous validation that the internal web application returns valid TLS certificates, responds within 500 ms, and does not exhibit unauthorized HTML tampering. The architecture forbids routing internal traffic across the public internet. How should the team configure CloudWatch Synthetics?
Deploy a CloudWatch Synthetics canary configured with the private VPC's subnets and security groups, ensuring the subnets have VPC endpoints for Amazon S3 and CloudWatch.
Configure a standard public CloudWatch Synthetics canary and whitelist AWS Synthetics public IP ranges on the internal ALB security group.
Create a Route 53 public health check pointing to the internal ALB private IP address using string-matching rules.
Deploy an EC2 instance running a cron job in each Availability Zone that executes curl commands and pushes custom metrics to CloudWatch.
A critical payment authorization daemon running on an EC2 instance emits a custom heartbeat metric named 'DaemonHeartbeat' to CloudWatch every 60 seconds with a value of 1. If an attacker gains root access and abruptly disables the monitoring agent or terminates the instance, the metric stream stops entirely. How should the security engineer configure the CloudWatch metric alarm so that it alerts the security team immediately upon cessation of the metric stream?
Set the evaluation statistic to SampleCount < 1 and configure 'Treat missing data as missing'.
Set the evaluation threshold to DaemonHeartbeat < 1 and configure 'Treat missing data as ignore'.
Set the evaluation threshold to DaemonHeartbeat == 0 and configure 'Treat missing data as notBreaching'.
Set the evaluation threshold to DaemonHeartbeat < 1 for 1 out of 1 datapoints and configure 'Treat missing data as breaching'.
An enterprise web portal uses Route 53 Active-Passive failover to route users to a secondary static error and disaster recovery bucket if the primary application fails. The application server was recently defaced by an attacker who modified the index page to display malicious content while the web server process continued returning HTTP status code 200 OK. How can the security team configure Route 53 health checking to detect such attacks and trigger automated failover?
Configure a Route 53 Calculated Health Check that averages TCP round-trip latency across all global edge locations.
Configure a Route 53 Endpoint Health Check with String Matching enabled, verifying the presence of a specific application integrity token within the first 5,120 bytes of the response body.
Switch the health check protocol from HTTPS to TCP port 443 to detect operating system compromise.
Attach a CloudWatch Synthetics canary to a Route 53 latency routing policy without health check association.
Sections you finish are checked off in the contents.