4.2 Threat Containment & Network Quarantine Techniques
Key Takeaways
Security group changes do not cut tracked connections, so quarantine an ENI by replacing all its groups with an allow-all group (making flows untracked) and then with a zero-rule isolation group; add a NACL deny for an instant, tracking-independent cutoff.
Stateless Network Access Control Lists (NACLs) provide immediate subnet-level isolation by placing a top-priority rule (Rule #1) explicitly denying all inbound and outbound traffic.
Revoking an IAM user or role's active STS temporary sessions requires applying an inline policy with an explicit Deny paired with the aws:TokenIssueTime condition, as deleting access keys only invalidates long-term credentials.
In Amazon EKS, compromised worker nodes must be cordoned rather than drained immediately, as draining evicts and destroys running container filesystems and volatile forensic memory.
AWS Backup Vault Lock in Compliance mode prevents unauthorized deletion, early expiration, or modification of recovery points, providing an immutable foundation for clean workload restoration.
4.2 Threat Containment & Network Quarantine Techniques
When a security incident is detected in an AWS environment, incident responders face an immediate operational dilemma: how to halt active malicious activity—such as data exfiltration, lateral movement, or ransomware encryption—without destroying volatile evidence or alerting the adversary to ongoing remediation. Containment is the bridge between detection and eradication. If executed clumsily, containment attempts can inadvertently alert an adversary (prompting them to execute destructive wiper malware) or destroy transient forensic data stored in RAM and network socket tables.
AWS provides granular control across network boundaries, identity planes, container runtimes, and immutable backup systems. Mastering these containment mechanisms is essential for any cloud security specialist.
Network Containment: Security Group Quarantine vs. Additive Traps
In Amazon EC2, virtual firewalls operate primarily at two layers: Security Groups (which act at the Elastic Network Interface / hypervisor layer) and Network Access Control Lists (NACLs) (which act at the subnet boundary).
Stateful Connection Tracking and Rule Swapping Mechanics
Security Groups are stateful. They utilize hypervisor connection tracking (conntrack) tables to automatically permit return traffic regardless of inbound or outbound rules. Understanding how security group rule changes interact with established connections is a frequent specialty exam focal point:
- Tracked connections survive rule changes: If you delete or narrow an allow rule, or swap the security groups on an ENI, established tracked connections are not interrupted; they continue until they time out. An attacker's existing reverse shell or C2 session can therefore survive a naive quarantine.
- Untracked flows are cut immediately: A flow is untracked when all traffic (0.0.0.0/0 or ::/0) is allowed in both directions and it is not an automatically tracked connection. Removing the rule that allows an untracked flow stops it at once. AWS isolation guidance uses this: first attach a security group that allows all inbound and outbound traffic (making existing flows untracked), then replace it with the isolation security group that has zero inbound and zero outbound rules.
- Instant, tracking-independent cutoff: Add a deny to the subnet NACL, which is stateless and evaluates every packet.
Warning
The Additive Security Group Trap: Security Groups can only define Allow rules; they cannot define Deny rules. Multiple security groups attached to a single network interface are evaluated with logical OR (additive permissive evaluation). If an instance has Web-SG (allowing TCP ports 80/443 inbound and all outbound) and a responder attaches Quarantine-SG (with zero rules) without detaching Web-SG, the instance remains 100% accessible! To achieve network quarantine, responders must replace all existing security groups with the Quarantine Security Group.
Designing the Quarantine Security Group
A robust quarantine security group architecture accommodates operational realities:
- Total Isolation Quarantine SG: Contains zero ingress and zero egress rules. Used when no remote interaction with the guest OS is required, or when volatile artifacts will be acquired via out-of-band methods (e.g., EBS snapshots and hypervisor-level metadata).
- Forensic Acquisition Quarantine SG: In scenarios where responders must acquire RAM using AWS Systems Manager Run Command, the quarantine SG must allow outbound HTTPS traffic (TCP port 443) strictly to the private IP CIDR ranges of the VPC Interface Endpoints for Systems Manager (
ssm,ssmmessages, andec2messages) and the S3 Gateway Endpoint. Zero inbound rules and zero internet egress rules are permitted. - Bastion Access Quarantine SG: Contains zero outbound rules, but contains a single inbound rule allowing SSH (port 22) or RDP (port 3389) exclusively from a hardened, isolated forensic analysis bastion host located in a private management subnet.
Stateless NACL Drop Rules & Subnet-Level Blast Radius
While Security Groups operate at the individual network interface level, Network Access Control Lists (NACLs) operate at the subnet boundary. NACLs are stateless, meaning inbound and outbound traffic must be explicitly permitted in both directions.
Ascending Numerical Evaluation Rule
NACLs process rules in ascending numerical order (e.g., 1 to 32766), stopping at the first rule that matches the packet headers (source IP, destination IP, protocol, port). At the end of every NACL sits an unmodifiable default rule (*) that denies all unmatched traffic.
NACL Inbound Rules Evaluation:
Rule 1: DENY TCP 0.0.0.0/0:ANY -> Matches packet? YES -> DROP IMMEDIATELY (Evaluation terminates)
Rule 100: ALLOW TCP 0.0.0.0/0:443 -> Never reached
Rule *: DENY ALL 0.0.0.0/0:ALL -> Default catch-all
Strategic Use Cases for NACLs During Containment
- Subnet-Level Quarantine: If an entire subnet is suspected of widespread worm propagation or automated lateral movement, editing each instance's security groups individually takes too long. Inserting Rule 1: DENY ALL 0.0.0.0/0 on both inbound and outbound NACL rules instantly severs all network traffic entering or leaving the entire subnet within milliseconds.
- Blocking Command-and-Control (C2) IP Ranges: When threat intelligence or GuardDuty identifies a specific external adversary IP or CIDR block (e.g.,
203.0.113.50/32), responders can insert a top-priority rule:Rule 10: DENY ALL 203.0.113.50/32. This blocks communication to and from the adversary across the entire subnet without interrupting legitimate application traffic. - Preserving Volatile Sockets vs. Dropping Packets: Because NACL drops occur at the subnet router boundary, the guest OS TCP stack remains alive. The OS does not send TCP RST packets unless the application times out, preserving connection state tables inside guest memory for forensic inspection.
Preserving Volatile Network States Before Containment
Before executing network isolation that clears connection tracking tables, responders should preserve active network connection telemetry if operational runbooks allow. Once network isolation is applied, active sockets transition to CLOSE_WAIT or TIME_WAIT and eventually disappear from kernel tables.
Volatile Network Artifacts Checklist
- Active Listening Ports and Sockets:
ss -tulpnornetstat -tulpn(identifying rogue daemons and reverse shells). - Active Established TCP Sessions:
ss -tapn(identifying external remote IP addresses, ephemeral ports, and associated Process IDs [PIDs]). - Kernel Connection Tracking Table:
conntrack -L(inspecting active NAT and tracked connection states). - Routing and ARP Cache Tables:
ip route showandip neigh show(identifying local network reconnaissance and spoofed gateways).
VPC Traffic Mirroring for Out-of-Band Capture
If an instance is actively communicating with an unknown external entity and responders need to inspect raw packet payloads without altering the host:
- Configure VPC Traffic Mirroring on the compromised instance's ENI.
- Define a Traffic Mirror Filter capturing all inbound and outbound protocols.
- Define a Traffic Mirror Target pointing to an Elastic Network Interface or Network Load Balancer attached to an out-of-band monitoring instance running Zeek, Suricata, or Wireshark.
- VPC Traffic Mirroring copies raw Layer 2 through Layer 7 packets directly at the Nitro hypervisor level, invisible to any rootkit or malware running inside the guest operating system.
Invalidating Active IAM Credentials & STS Temporary Sessions
When credentials are leaked or an identity is compromised, revoking access requires understanding the fundamental difference between long-term IAM user credentials and temporary AWS Security Token Service (STS) credentials.
The Long-Term vs. Short-Term Credential Trap
If an IAM user's secret access key is exposed in a public Git repository, an incident responder might immediately deactivate or delete the access key using iam:UpdateAccessKey --status Inactive.
Caution
The STS Session Persistence Trap: Deactivating or deleting an IAM access key only blocks future requests that authenticate directly with that long-term key. If the adversary already used that access key to call sts:GetSessionToken or sts:AssumeRole, the resulting temporary credentials (ASIA...) remain fully active and functional until their configured session duration expires (up to 36 hours for GetSessionToken or GetFederationToken sessions, or up to 12 hours for assumed roles)! Deleting the IAM user or access key does NOT revoke previously minted STS temporary tokens.
The aws:TokenIssueTime Revocation Mechanism
To immediately invalidate active temporary credentials issued by AWS STS, security teams must attach an inline policy to the compromised IAM role or user containing an explicit Deny paired with the aws:TokenIssueTime condition key.
Temporary credentials issued by STS include a cryptographic timestamp indicating when the token was generated (TokenIssueTime). An IAM policy condition checking DateLessThan against aws:TokenIssueTime forces the IAM evaluation engine to reject any request signed by a token minted prior to the specified cutoff time:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "RevokeAllTemporarySessionsBeforeRevocationTime",
"Effect": "Deny",
"Action": "*",
"Resource": "*",
"Condition": {
"DateLessThan": {
"aws:TokenIssueTime": "2026-09-29T15:30:00Z"
}
}
}
]
}
Why This Works
Under the AWS Policy Evaluation Logic:
- Explicit Deny overrides all Allow permissions.
- Any API call presented with a temporary token issued before
2026-09-29T15:30:00Zevaluates the condition asTrue, triggering the explicit Deny and returning anAccessDeniederror. - Legitimate services or users can subsequently re-authenticate or re-assume the role; tokens issued after the revocation timestamp will evaluate the condition as
False, bypassing the Deny and restoring normal access.
Exam Tip: In the IAM Management Console, clicking the "Revoke active sessions" button on an IAM role automatically attaches an inline policy named
AWSRevokeOlderSessionsthat implements this exactaws:TokenIssueTimepattern, using a timestamp about 30 seconds in the future to allow for propagation. For an IAM user, attach an equivalent deny policy yourself.
Container Incident Isolation in Amazon EKS
When a container running inside Amazon Elastic Kubernetes Service (Amazon EKS) is compromised (e.g., via a remote code execution vulnerability or a compromised container registry image), containment requires isolating the workload at both the Kubernetes control plane and the AWS network layers.
Node Cordoning vs. Pod Eviction / Draining
When a container escape or malicious pod is identified on an EKS worker node, administrators must avoid reflexive eviction:
| Command | Operational Action | Forensic Consequence |
|---|---|---|
kubectl cordon <node> | Marks the worker node as Unschedulable. Existing running pods remain untouched; no new pods can be placed on the node. | Recommended for forensics. Preserves running container namespaces, volatile memory, and local ephemeral filesystems for analysis. |
kubectl drain <node> | Evicts all running pods from the node, terminating container processes and rescheduling them on other worker nodes. | Forensic Anti-Pattern. Evicting a pod issues a SIGTERM / SIGKILL to container processes, immediately deleting container memory and wiping ephemeral writeable container layers (overlay2). |
kubectl delete pod <pod> | Deletes the pod object immediately. | Wipes container runtime state; if managed by a Deployment, a new identical (potentially vulnerable) pod immediately launches. |
Isolating Pods with Kubernetes NetworkPolicies
If the EKS cluster uses a CNI plugin supporting Kubernetes NetworkPolicies (such as Calico, Cilium, or the Amazon VPC CNI with Network Policy support enabled), responders apply an explicit default-deny NetworkPolicy matching the labels of the compromised pod:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: quarantine-compromised-workload
namespace: payment-processing
spec:
podSelector:
matchLabels:
app: invoice-generator
policyTypes:
- Ingress
- Egress
By specifying policyTypes: ["Ingress", "Egress"] without providing any ingress or egress rule blocks, the NetworkPolicy drops 100% of network traffic into and out of matching pods while leaving the container processes running in memory for live forensic triage.
Security Groups for Pods (Amazon VPC CNI)
In enterprise EKS environments utilizing Security Groups for Pods, individual Kubernetes pods receive dedicated Elastic Network Interfaces (branch ENIs). In this architecture, responders can modify the pod's underlying security group directly via the AWS EC2 API, attaching an isolated Quarantine Security Group without modifying Kubernetes cluster manifests.
Eradication & Clean Restoration via AWS Backup
Once an incident is contained and forensic artifacts are secured, the Eradication and Recovery phase begins. In modern cloud architecture, trying to manually locate and remove rootkits, backdoors, or malicious cron jobs from a compromised operating system is considered an anti-pattern.
The "Cattle, Not Pets" Eradication Principle
- Terminate Compromised Compute: Never sanitize an EC2 instance in production. Once evidence is captured, terminate the instance (
ec2:TerminateInstances). - Redeploy from Golden AMIs: Launch clean, verified instances from immutable Golden AMIs or container base images created via hardened CI/CD pipelines (e.g., EC2 Image Builder).
- Rotate All Underlying Secrets: Before launching replacement workloads, rotate all database credentials, API tokens, and private keys stored in AWS Secrets Manager and AWS KMS.
Restoring Data from Immutable Backups: AWS Backup Vault Lock
For persistent data stores, relational databases (Amazon RDS), and file systems (Amazon EFS), recovery must draw from verified, uncorrupted backup points.
However, sophisticated adversaries (such as ransomware operators) frequently target backup repositories before executing their primary payload, attempting to delete recovery points to prevent organizational recovery.
AWS Backup Vault Lock protects recovery points against unauthorized deletion and retention changes:
- Governance Mode: Allows privileged IAM principals with dedicated permissions to delete recovery points or shorten retention windows. Suitable for testing.
- Compliance Mode: Enforces strict, irreversible WORM storage. Once the configurable grace period (a cooling-off window of at least 3 days, or 72 hours) expires, no entity—including the AWS account root user or AWS Support—can delete backups or reduce retention periods until the recovery points naturally expire.
- Cross-account copies and logically air-gapped vaults: Copy recovery points to a dedicated Backup account. AWS Backup's logically air-gapped vault goes further: it is always locked in compliance mode, stores backups in an AWS Backup service-owned account, is encrypted with an AWS owned key by default (or a customer managed key), and can be shared through AWS RAM or recovered through multi-party approval, so you can restore even if the owning account is compromised.
A production Amazon EC2 instance hosting a customer-facing API is detected communicating with a cryptomining mining pool. The security operations team must immediately isolate the instance from the network to prevent further outbound connections while keeping the instance running to facilitate live forensic memory capture. The instance currently has two security groups attached: 'Web-Tier-SG' and 'SSH-Admin-SG'. How should the security team configure the instance's security groups to achieve immediate isolation?
Create a new security group named 'Quarantine-SG' containing an explicit Deny rule for all outbound traffic, and attach it alongside 'Web-Tier-SG' and 'SSH-Admin-SG'.
Attach a new security group named 'Quarantine-SG' containing zero inbound and zero outbound rules to the instance, leaving the existing security groups attached.
Replace both 'Web-Tier-SG' and 'SSH-Admin-SG' with an allow-all security group so existing flows become untracked, then replace that with a single 'Quarantine-SG' containing zero inbound and zero outbound rules.
Edit 'Web-Tier-SG' and 'SSH-Admin-SG' to remove their inbound rules, while leaving the default outbound allow rule (0.0.0.0/0) active to allow security log transmission.
A rogue insider obtained long-term IAM user access keys for a DevOps administrator and used them to assume an administrative IAM role ('ProductionDeployerRole') via AWS Security Token Service (STS). The insider is actively launching unauthorized EC2 instances across multiple AWS Regions. The security team immediately deletes the IAM user's access keys, but notices that unauthorized API calls continue using the assumed role session. How can the security team immediately terminate the insider's ability to execute API calls using the assumed role?
Attach an inline policy to 'ProductionDeployerRole' containing an explicit Deny for all actions with a condition checking that aws:TokenIssueTime is DateLessThan the current UTC timestamp.
Delete 'ProductionDeployerRole' from the IAM console and recreate it with the exact same name and trust policy.
Modify the IAM role's maximum session duration setting from 12 hours to 1 hour in the role configuration.
Call the sts:DecodeAuthorizationMessage API to invalidate the active session token cache across AWS STS endpoints.
A security alert indicates that a container pod inside an Amazon EKS cluster has been compromised by an attacker executing an interactive reverse shell. The security engineer needs to contain the compromised workload to prevent lateral network reconnaissance while preserving the container's volatile memory and filesystem state for an upcoming forensic investigation. What action should the engineer take first?
Execute kubectl drain on the Kubernetes worker node hosting the compromised pod to evict and isolate all workloads.
Execute kubectl delete pod on the compromised pod to force the deployment controller to spin up a clean replica.
Terminate the underlying EC2 worker node via the Amazon EC2 console to force Kubernetes to reschedule pods onto healthy nodes.
Execute kubectl cordon on the worker node to prevent new pod scheduling, and apply a Kubernetes NetworkPolicy to the pod specifying empty ingress and egress policy types.
An enterprise organization was recently subjected to a ransomware incident where an attacker gained administrative credentials and attempted to wipe all production data and database backups. The security team needs to implement an immutable backup architecture using AWS Backup to ensure that recovery points for Amazon RDS databases and Amazon EBS volumes cannot be deleted, altered, or have their retention periods shortened by any administrative entity, including compromised root user credentials. Which solution meets these requirements?
Store all backups in an S3 bucket configured with S3 Versioning and an S3 Lifecycle rule that transitions objects to S3 Glacier Flexible Retrieval.
Deploy an AWS Backup vault configured with AWS Backup Vault Lock in Compliance mode with appropriate minimum and maximum retention periods.
Configure AWS Backup Vault Lock in Governance mode and assign the backup vault deletion permission exclusively to a break-glass IAM role.
Deploy an AWS Lambda function triggered by Amazon EventBridge on backup:DeleteRecoveryPoint API calls to automatically recreate deleted snapshots.
Sections you finish are checked off in the contents.