4.3 Correlating Events & Root Cause Analysis with Amazon Detective
Key Takeaways
Amazon Detective automatically builds and continuously updates a graph-based data model (Behavior Graph) using machine learning and graph algorithms, retaining up to 12 months of security telemetry.
Detective's core sources are AWS CloudTrail logs, Amazon VPC Flow Logs, and Amazon GuardDuty findings; EKS audit logs and AWS security findings (from Security Hub CSPM) are optional source packages, all ingested without customer-built pipelines.
Finding Groups correlate disparate GuardDuty findings across multiple resources and accounts into unified attack chains, highlighting the full blast radius and root cause.
Detective provides behavioral baselining to visualize anomalous API volumes, newly assumed roles, and unusual geographic locations compared against historical norms.
Root cause analysis identifies initial compromise vectors—including exposed IAM credentials, Server-Side Request Forgery (SSRF) targeting IMDSv1, and unpatched web vulnerabilities—allowing teams to implement definitive architectural defenses such as mandatory IMDSv2.
4.3 Correlating Events & Root Cause Analysis with Amazon Detective
Identifying that an incident has occurred is merely the opening chapter of incident response. Once compromised resources are contained, security teams must answer critical investigative questions: How did the adversary gain initial access? What identities and resources were touched? Did the attacker pivot laterally across accounts? What data was exfiltrated?
In complex multi-account AWS Organizations, answering these questions manually is daunting. Telemetry is scattered across billions of raw log records: AWS CloudTrail management events, CloudTrail S3 data events, Amazon VPC Flow Logs, Route 53 DNS query logs, and Amazon EKS audit logs. Security analysts attempting to query these logs via Amazon Athena or OpenSearch often struggle with query syntax, missing cross-account context, and the sheer computational latency of querying terabytes of unindexed data.
Amazon Detective solves this challenge by transforming raw security telemetry into an interactive, graph-based data model known as the Behavior Graph, enabling security responders to conduct root cause analysis and blast radius mapping in minutes rather than days.
Amazon Detective Architecture & The Behavior Graph
Amazon Detective is a fully managed security investigation service that automatically ingests, structures, and correlates security data from multiple AWS data sources. Rather than functioning as a standard log search engine, Detective uses graph modeling, statistical analysis, and machine learning to construct a dynamic, interconnected representation of your cloud environment.
Graph Primitives: Nodes and Edges
The Behavior Graph organizes telemetry into two primary elements:
- Nodes (Entities): Represent distinct AWS resources and external actors. Entities include AWS Accounts, IAM Users, IAM Roles, Federated User Sessions, Amazon EC2 Instances, Amazon S3 Buckets, IP Addresses, and User Agents.
- Edges (Relationships): Represent interactions between entities observed over time. Relationships include API calls executed by an IAM role, network connections established between an EC2 instance and an external IP address, roles assumed by an identity, and Kubernetes pods deployed by a specific service account.
Historical Baselines (12-Month Rolling Window)
A critical advantage of Amazon Detective is its 12-month rolling retention window. Detective establishes behavioral baselines for each entity over time, allowing analysts to compare current suspicious activity against historical norms:
- Is it normal for this IAM role to invoke
iam:CreateAccessKey? - Has this EC2 instance ever communicated with this external IP address or Autonomous System Number (ASN) in the past 180 days?
- Does this administrative user typically authenticate from this geographic country or ISP?
Zero-ETL Ingestion Architecture
A major architectural highlight tested on the exam is Detective's serverless, out-of-band ingestion:
- No Customer Logging Pipelines: You do not need to configure Amazon S3 export buckets, Amazon Kinesis streams, or custom CloudWatch log subscription filters to feed Detective.
- No Performance Impact: Telemetry is ingested directly from the AWS service backplanes out-of-band, introducing zero latency or CPU overhead to running EC2 instances or container workloads.
- No Duplicate Ingestion Charges: Ingesting VPC Flow Logs or CloudTrail into Detective does not incur CloudWatch Logs ingestion fees or duplicate S3 storage fees for Detective's internal graph processing.
Multi-Account Governance with AWS Organizations
Amazon Detective integrates seamlessly with AWS Organizations:
- The Organizations Management Account designates a specific account (typically the central Security Tooling Account) as the Detective Delegated Administrator.
- The Delegated Administrator can automatically enable Detective across all existing member accounts and configure auto-enablement for any newly created AWS accounts in the organization.
- Telemetry from all member accounts is synthesized into a single, unified enterprise Behavior Graph, allowing cross-account pivot analysis from a centralized pane of glass.
Finding Groups and Visual Blast Radius Analysis
During a multi-stage security intrusion, attackers rarely generate a single, isolated alert. An adversary might perform port scanning, compromise an EC2 web server, harvest temporary IAM credentials via metadata services, enumerate S3 buckets, and exfiltrate sensitive customer databases. In standard security monitoring, this sequence generates half a dozen separate Amazon GuardDuty findings across multiple hours or days.
Amazon Detective Finding Groups
To prevent alert fatigue and surface the broader attack campaign, Amazon Detective automatically clusters related GuardDuty findings into Finding Groups.
Finding Groups use graph connectivity algorithms to correlate findings that share common entities (such as the same IP address, IAM user, EC2 instance, or VPC). Instead of triaging disconnected alerts, security analysts open a single Finding Group that displays:
- The complete attack timeline from initial reconnaissance to objective completion.
- The root entity that served as the entry point.
- All interconnected resources that have been touched or compromised.
Visual Blast Radius Analysis
When triaging a compromised entity, Detective provides interactive visual graph panels that map its blast radius:
| Investigative View | Telemetry Analyzed | Critical Investigative Insight |
|---|---|---|
| Resource Interaction Graph | CloudTrail logs | Visualizes all AWS services and S3 buckets accessed by an IAM role during the incident window. |
| Overall Network Activity | Amazon VPC Flow Logs | Displays inbound and outbound traffic volume (in bytes and packets), broken down by protocol, port, and direction. |
| New / Unfamiliar Connections | Historical VPC Flow Baselines | Highlights external IP addresses and geolocations that have never communicated with the instance prior to the incident. |
| EKS Pod Execution & Audit | Kubernetes Audit Logs | Shows API requests executed inside EKS clusters, mapping pod creations, privileged escalations, and namespace access. |
Investigating Behavioral Anomalies: API Volumes, Geolocation & Roles
Amazon Detective visually plots observed entity behavior against established historical baselines, allowing analysts to rapidly spot deviations that indicate compromise.
1. Anomalous API Volume Spikes
When an adversary compromises IAM credentials, they frequently execute automated discovery scripts (e.g., Pacu, ScoutSuite) or initiate mass data exfiltration. Detective tracks daily and hourly API call volumes for every IAM identity, comparing observed volume against a 45-day rolling average.
In the Detective console, API activity is plotted on an interactive timeline:
- Green Baseline: Expected normal API volume for that identity.
- Blue Line: Actual observed API volume.
- Red Anomalies: Statistical breaches where call volume exceeds expected prediction bands by multiple standard deviations.
Responders can click directly on an anomalous volume spike to filter down to the exact API operations responsible (e.g., a surge in s3:GetObject or secretsmanager:GetSecretValue calls).
2. Geolocation and IP Address Context
CloudTrail records the source IP address (sourceIPAddress) for every API call, but IP addresses alone lack operational context. Detective enriches every IP address with:
- Autonomous System Number (ASN) and ISP Details: Differentiates residential ISPs, corporate VPNs, and public cloud hosting providers (e.g., DigitalOcean, Linode, AWS).
- Geolocation Mapping: Identifies the geographic country and city of origin.
- Historical Familiarity: Displays whether the IP address or geographic location is New to Account or New to Role. If an administrative role that has operated exclusively from North America for 12 months suddenly executes API calls from an unmapped foreign ASN, Detective highlights this as an immediate anomaly.
- Multiple Roles from Single External IP: Detective's IP address profile shows when multiple distinct IAM roles or users across different AWS accounts are accessed from the exact same external IP address—a hallmark indicator that an external attacker is testing multiple stolen credentials.
3. Newly Assumed Roles and Cross-Account Pivots
Adversaries who compromise an initial IAM role frequently attempt privilege escalation by assuming secondary roles (sts:AssumeRole) across accounts. Detective's Behavior Graph tracks role assumption chains:
- Identifies when an identity assumes a role it has never historically assumed.
- Maps cross-account role assumption paths, allowing investigators to follow an attacker's movement from a development account into a production database account.
Identifying the Initial Compromise Vector: Deep-Dive Scenarios
Root cause analysis requires tracing an attack backward through the Behavior Graph to identify how the adversary obtained access. In the SCS-C03 exam, three primary initial compromise vectors dominate:
Scenario 1: Exposed Long-Term IAM Access Keys
- The Attack: An engineer inadvertently commits an IAM user's long-term access key (
AKIA...) to a public GitHub repository. Automated botnets scrape the key within minutes. - Detective Evidence Trail:
- Detective shows the IAM user entity executing API calls from an external IP belonging to a known commercial hosting provider or Tor exit node.
- Geolocation shows "New to User" and "New to Account".
- CloudTrail management events reveal rapid reconnaissance calls:
iam:GetAccountSummary,iam:ListUsers,ec2:DescribeInstances. - The user agent string identifies automated command-line tooling (e.g.,
aws-sdk-go/v1.38,python-requests).
- Architectural Remediation:
- Revoke active STS sessions via inline policy with
aws:TokenIssueTimecondition. - Delete the compromised access key (
iam:DeleteAccessKey). - Implement AWS Secrets Manager for programmatic secret storage.
- Deploy IAM Access Analyzer external access findings and automated repository scanning (e.g., GitHub secret scanning).
- Revoke active STS sessions via inline policy with
Scenario 2: Server-Side Request Forgery (SSRF) and IMDSv1 Credential Exfiltration
- The Attack: An attacker discovers a Server-Side Request Forgery (SSRF) vulnerability in an EC2-hosted web application. The attacker crafts an HTTP request forcing the web server to query the local Instance Metadata Service (IMDS) endpoint at
http://169.254.169.254/latest/meta-data/iam/security-credentials/<RoleName>. Because the instance uses IMDSv1, which supports plain HTTP GET requests without session authentication, the instance metadata service returns temporary role credentials (ASIA..., SecretAccessKey, and Token) in plain text to the attacker. The attacker copies these credentials to their external laptop and makes API calls directly to AWS.
- Detective Evidence Trail:
- GuardDuty fires the finding
UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration. - In Detective, the analyst inspects the finding group. Detective displays the EC2 instance role node connected to an external, untrusted IP address that has never been associated with the VPC.
- Detective's resource interaction graph shows the credentials being used to call
s3:ListBucketsands3:GetObjectfrom the external IP, proving credential exfiltration.
- GuardDuty fires the finding
- Architectural Remediation:
- Enforce IMDSv2: IMDSv2 is session-oriented. It requires a client to obtain a cryptographic session token via an HTTP
PUTrequest with the headerX-aws-ec2-metadata-token-ttl-seconds: 21600before accessing metadata. Web application SSRF vulnerabilities almost universally cannot execute arbitrary HTTP PUT requests with custom headers. - Restrict the HTTP Put Response Hop Limit to 1 (
--http-put-response-hop-limit 1). This ensures that even if a container running on the EC2 host attempts an SSRF, the metadata response packet cannot traverse an IP forwarding boundary (bridge or container network). - Enforce IMDSv2 across the organization using an AWS Organizations Service Control Policy (SCP):
- Enforce IMDSv2: IMDSv2 is session-oriented. It requires a client to obtain a cryptographic session token via an HTTP
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "EnforceIMDSv2AcrossOrganization",
"Effect": "Deny",
"Action": "ec2:RunInstances",
"Resource": "arn:aws:ec2:*:*:instance/*",
"Condition": {
"StringNotEquals": {
"ec2:MetadataHttpTokens": "required"
}
}
}
]
}
Scenario 3: Unpatched Web Vulnerability & Remote Code Execution (RCE)
- The Attack: An internet-facing Apache or Java application server is compromised via an unpatched Remote Code Execution vulnerability (e.g., Log4Shell). The adversary executes shell commands, spawns a reverse shell, and downloads an encrypted cryptominer binary.
- Detective Evidence Trail:
- GuardDuty Runtime Monitoring (eBPF agent) detects the malicious execution (
Execution:Runtime/NewBinaryExecutedorCryptoCurrency:Runtime/BitcoinTool.B). - In Detective, the analyst examines the EC2 instance node's Overall Network Activity.
- A massive, sudden spike in outbound network traffic appears on non-standard ports (e.g., TCP port 3333 or 8333).
- Detective enriches the destination external IP address, identifying it as a known mining pool.
- GuardDuty Runtime Monitoring (eBPF agent) detects the malicious execution (
- Architectural Remediation:
- Network containment: swap attached Security Groups with a quarantine SG.
- Eradicate the compromised instance and redeploy from a patched Golden AMI.
- Implement AWS WAF with AWS Managed Rules (e.g.,
AWSManagedRulesKnownBadInputsRuleSetandAWSManagedRulesCommonRuleSet) in front of the Application Load Balancer to inspect and block malicious payload headers. - Automate vulnerability scanning using Amazon Inspector with automated remediation workflows.
Validating Findings: Scope and Impact (Skill 2.2.3)
SCS-C03 added an explicit skill for validating findings from AWS security services to assess the scope and impact of an event. Before containment starts, answer three questions with evidence:
- Is the finding a true positive? Compare the finding's principal, resource, and activity with expected behavior. A GuardDuty reconnaissance finding caused by the company's own authorized vulnerability scanner is expected activity; handle it with a narrow suppression rule, not by ignoring the finding type. Check CloudTrail for the underlying API calls: calls that failed with
AccessDeniedindicate an attempt, while successful calls indicate impact. - What is the scope? Identify every account, Region, principal, and resource involved. Detective finding groups and entity profiles show related findings and entities; GuardDuty Extended Threat Detection correlates multi-stage activity into attack sequence findings; Security Hub CSPM cross-Region aggregation shows the same resource across findings; and an Athena or CloudTrail Lake query on the compromised
userIdentity.accessKeyIdshows every call that key made in every account. - What is the impact? Determine what data or configuration was touched. S3 data events in CloudTrail show which objects were read; Amazon Macie shows whether those buckets hold sensitive data; IAM Access Analyzer and the role's policies show what the principal could have reached; VPC Flow Logs byte counts indicate possible exfiltration volume; and Amazon Inspector findings show whether affected instances had exploitable vulnerabilities.
| Question | Primary Evidence |
|---|---|
| Did the attacker's calls succeed? | CloudTrail errorCode and errorMessage fields |
| Which other resources and accounts are involved? | Detective finding groups, GuardDuty attack sequences, organization-wide CloudTrail queries |
| Was sensitive data exposed? | S3 data events plus Macie classification of the affected buckets |
| How much data left the network? | VPC Flow Logs bytes to external addresses |
| Could the principal escalate further? | IAM policies, permissions boundaries, and Access Analyzer findings |
Record the validated scope, impact, and any severity change in the incident record (for example, an OpsCenter OpsItem) so that containment targets every affected resource and the post-incident review has an evidence trail.
Specialty Exam Pitfalls & Traps
- Assuming Detective Requires S3/CloudWatch Ingestion Setup: Unlike Amazon Athena or custom SIEM solutions, Detective does not require configuring S3 export buckets, Kinesis firehoses, or CloudWatch log groups. Ingestion is fully handled natively and out-of-band by AWS.
- Confusing GuardDuty with Detective: GuardDuty is a detection engine that monitors telemetry and emits findings when threats are identified. Detective is an investigation tool that takes those findings, correlates them with multi-account telemetry, and builds a visual behavior graph for root cause analysis.
- Overlooking IMDSv2 Hop Limit: Mandating
ec2:MetadataHttpTokens: requiredis only half the battle. If workloads run inside containers (Docker / EKS) on the EC2 host, the hop limit must be explicitly set to1so container processes cannot forward tokens off-host. - Single-Finding Triage vs. Finding Groups: When investigating complex incidents, do not analyze GuardDuty findings in isolation. Use Detective Finding Groups to understand the full multi-stage kill chain and blast radius.
A security operations team is investigating a complex security breach involving multiple AWS accounts in an AWS Organization. The team needs to correlate Amazon GuardDuty findings, AWS CloudTrail management events, VPC Flow Logs, and Amazon EKS audit logs across all accounts to determine the initial compromise vector and visual blast radius. The solution must minimize operational maintenance and require zero custom log ingestion pipelines or database management. Which architecture should the team deploy?
Configure Amazon EventBridge rules in each member account to stream all logs to an Amazon Data Firehose in a central logging account, load data into Amazon OpenSearch Service, and build custom correlation dashboards.
Deploy an Amazon Athena federated query architecture that executes scheduled SQL joins across S3 log buckets in every member account, exporting results to Amazon QuickSight for visualization.
Enable AWS Security Hub across all accounts, export all ASFF findings to an Amazon Aurora PostgreSQL database, and use an AWS Lambda function to calculate graph relationships.
Designate the central security tooling account as the Amazon Detective delegated administrator in AWS Organizations, enable Detective across all member accounts, and use Detective Finding Groups and the Behavior Graph for investigation.
A security analyst is investigating an Amazon GuardDuty finding: 'UnauthorizedAccess:IAMUser/InstanceCredentialExfiltration'. The finding indicates that temporary security credentials assigned to an EC2 instance's IAM role ('AppServerRole') were used to execute API calls from an external IP address located outside AWS IP ranges. Amazon Detective confirms that the instance metadata service on the web server was accessed via a Server-Side Request Forgery (SSRF) vulnerability. Which architectural remediation should the security team implement to definitively eliminate this attack vector across all EC2 workloads?
Attach an inline IAM policy to 'AppServerRole' denying ec2:DescribeInstances calls from non-VPC IP CIDR ranges.
Enforce IMDSv2 on all EC2 instances by configuring MetadataHttpTokens to 'required' and setting the HTTP Put Response Hop Limit to 1, enforced via an AWS Organizations Service Control Policy.
Deploy an AWS Network Firewall rule that blocks outbound TCP port 80 traffic originating from the web server subnet.
Switch the EC2 instance tenancy from shared to dedicated host tenancy to prevent hypervisor-level memory inspection.
During an ongoing incident investigation, a security analyst notices multiple high-severity Amazon GuardDuty alerts firing across different member accounts: reconnaissance port scans on an EC2 instance, followed by IAM credential enumeration, and an anomalous volume of S3 data exfiltration. The alerts were emitted over a 72-hour window. How does Amazon Detective assist the analyst in understanding the unified scope of this incident rather than analyzing each alert in isolation?
Detective automatically combines related GuardDuty findings that share common resources, IP addresses, or entities into Finding Groups, presenting a single visual attack graph and unified timeline.
Detective automatically terminates the EC2 instances and deletes the S3 buckets associated with any finding whose severity score exceeds 7.0.
Detective exports the raw VPC Flow Logs and CloudTrail events into a PDF incident report sent via Amazon SES to the security team.
Detective generates AWS WAF rate-limiting rules and deploys them to all Application Load Balancers in the affected accounts.
A financial enterprise uses Amazon Detective across its AWS Organization. An incident responder investigates an alert involving an IAM user who allegedly created unauthorized IAM access keys. While reviewing the IAM user entity in Amazon Detective, the responder notices that the API call volume panel displays a significant red anomaly. What baseline comparison does Amazon Detective use to calculate and display this anomalous activity?
Detective compares the user's API call volume against the static thresholds configured in AWS CloudTrail event selectors.
Detective compares the user's API call volume against the industry average call volumes published in AWS Security Hub benchmarks.
Detective compares the user's observed API call volume against that specific identity's 45-day historical baseline of normal activity using statistical prediction bands.
Detective compares the user's API call volume against the root user's maximum permitted service quota in that specific AWS Region.
Sections you finish are checked off in the contents.