9.3 Cloud Incident Containment, Snapshot Forensics, and Container Security

Key Takeaways

  • A quarantine security group blocks new AWS EC2 flows, but tracked connections may persist; immediate isolation requires a tested stateless control such as a network ACL while responders preserve volatile evidence and remove cloud permissions.

  • Revoking long-lived IAM credentials does not invalidate previously issued STS tokens; use the role's Revoke active sessions action or an equivalent explicit Deny based on aws:TokenIssueTime, then account for IAM policy propagation and verify access is denied.

  • Cloud snapshots of running disks are generally crash-consistent rather than automatically application-consistent; preserve provenance, copy evidence to an isolated account, and use filesystem-specific no-recovery read-only mounts.

  • Container escapes exploit shared Linux kernels, dangerous capabilities (CAP_SYS_ADMIN), and mounted Docker sockets (/var/run/docker.sock) to compromise underlying container host nodes.

  • Ephemeral-container triage should prevent rescheduling, preserve runtime metadata and writable-layer changes, and acquire supported volatile evidence before the orchestrator destroys the workload; pausing alone is not a memory capture.

Last updated: October 2026

Cloud Incident Containment, Snapshot Forensics, and Container Security

Containing security incidents in public cloud environments presents fundamentally different operational requirements compared to traditional on-premises forensics. In cloud environments, abruptly terminating a compromised virtual machine (AWS EC2, Azure VM, GCP Compute Engine) obliterates volatile memory (RAM), wipes ephemeral storage, and triggers auto-scaling groups to provision replacement instances, perpetuating the intrusion. Cloud containment and forensic investigation rely on software-defined network isolation, non-destructive snapshot acquisition, and specialized workflows for serverless and ephemeral container architectures.

Cloud Containment Playbooks: Virtual Machines and Identities

Incident responders execute coordinated containment playbooks across both virtual compute workloads and cloud identity services:

  • Host Isolation via Layered Network Controls: Rather than powering down a compromised virtual machine, responders can attach a dedicated quarantine security group or Azure NSG that permits only approved forensic traffic. This blocks new flows while the machine remains running for volatile triage. AWS security groups are stateful, however, and changing a rule does not necessarily interrupt tracked connections immediately. If an active C2 channel must be cut at once, responders need a tested stateless control—such as an approved deny rule in the subnet network ACL—or another provider-supported isolation action, with scope and business impact assessed because a subnet ACL can affect other workloads.
  • Detaching IAM Instance Profiles: Compromised virtual machines often execute with attached IAM roles (EC2 Instance Profiles or Azure Managed Identities) granting access to cloud resources. Threat actors exploit these roles to issue API calls against the cloud control plane. Handlers must immediately detach the instance profile or replace it with an empty IAM role containing an explicit Deny policy, preventing the host from accessing cloud storage or pivoting across accounts.
  • Identity Revocation and Session Invalidation: Revoking compromised credentials requires addressing both static and temporary tokens:
    • Long-Lived Access Keys: For compromised IAM users, responders immediately change access key status to Inactive or delete the key, reset console passwords, and terminate active web sessions.
    • Temporary STS Credentials: Deleting an IAM access key does not invalidate temporary security tokens previously generated through sts:AssumeRole or federation. These tokens remain valid until their expiration timestamp. To revoke permissions from earlier role sessions, incident handlers can use the role's Revoke active sessions action or attach an equivalent inline policy containing an explicit Deny statement with an aws:TokenIssueTime condition (e.g., "Condition": {"DateLessThan": {"aws:TokenIssueTime": "2026-10-03T20:00:00Z"}}). After IAM policy propagation, this causes authorization to deny requests made with role sessions issued before the cutoff. Responders must verify the policy scope and effect rather than assume instantaneous global enforcement.

Forensically Sound Cloud Evidence Acquisition

Cloud infrastructure provides robust, API-driven evidence acquisition mechanisms that preserve chain of custody without physical media handling:

  • Point-in-Time Disk Snapshotting: Cloud providers offer block-level snapshots (AWS EBS, Azure Managed Disk, and GCP Persistent Disk). A snapshot of a running volume is generally crash-consistent, not automatically application-consistent: cached writes and multi-volume transactions may be incomplete unless responders can safely quiesce the workload. Record the API request, account, region, volume identifiers, timestamps, and encryption context; copy the snapshot into an isolated account; and validate exported or mounted artifacts with cryptographic hashes. A snapshot preserves storage blocks, not volatile RAM.
  • Cross-Account Isolation and Chain of Custody: If an adversary compromises the production cloud account, snapshots stored within that same account remain vulnerable to tampering or deletion. To ensure evidence integrity and establish a defensible chain of custody, incident handlers share the snapshot with a dedicated, isolated Forensic Investigation Account and copy it into that account. During the cross-account copy, the snapshot is re-encrypted using a new, isolated Customer Master Key (CMK) managed by AWS KMS accessible exclusively by forensic personnel.
  • Forensic Workstation Analysis: In the isolated forensic account, responders create an EBS or managed disk volume from the secured snapshot and attach it as a secondary, non-boot data disk to an analysis instance (such as a SIFT Workstation). Crucially, attach the derived analysis volume as a secondary disk and prevent journal replay. For ext4, an example is mount -o ro,noload /dev/xvdf1 /mnt/evidence; XFS uses filesystem-specific options such as ro,norecovery. Work from a copied snapshot, document every transformation, and hash exported evidence. Read-only mounting reduces modification risk but does not by itself prove forensic integrity.

Serverless and Container Incident Response

Modern cloud-native workloads introduce distinct containment and forensic challenges:

  • Serverless Architectures (AWS Lambda, Azure Functions): Serverless computing abstracts virtual infrastructure entirely. Functions execute within short-lived, stateless microVM containers terminating within seconds or minutes. Responders cannot dump physical memory or snapshot virtual disks. Incident response relies on centralized logging (CloudWatch Logs, Azure Monitor), tracking input payloads (API Gateway triggers), inspecting environment variables for leaked credentials, and examining temporary storage directories (such as /tmp in Lambda) where attackers stage payloads during execution.
  • Container Threat Vectors and Escapes: Containerized environments (Docker and Kubernetes) introduce unique attack surfaces. Attackers exploit malicious public images, misconfigured Kubernetes RBAC roles, and vulnerable application dependencies. The most dangerous container threats involve Container Escapes, where an adversary breaks out of containerized isolation to gain host-level operating system control. Primary escape vectors include:
    • Exploiting shared Linux kernel vulnerabilities (e.g., Dirty COW, Dirty Pipe).
    • Exposed Docker Sockets: Containers configured with mounted host Docker sockets (-v /var/run/docker.sock:/var/run/docker.sock) enable an adversary inside the container to communicate directly with the host Docker daemon, creating privileged containers that compromise the entire host.
    • Excessive Linux Capabilities: Containers executed with the --privileged flag or granted dangerous capabilities like CAP_SYS_ADMIN or CAP_NET_ADMIN can manipulate cgroups, load malicious kernel modules, or bypass AppArmor and SELinux profiles.
  • Ephemeral Container Forensic Triage: Because Kubernetes automatically terminates and replaces failing containers (CrashLoopBackOff), forensic evidence can vanish in seconds. When responding to container incidents, handlers must:
    1. Cordon the host node to prevent new pods from scheduling and, where the runtime and safety plan permit, pause the suspect container to stop execution while collection begins. Pausing does not create a memory image; responders still need a supported memory-acquisition method before the node or workload is destroyed.
    2. Capture live file modifications by running docker diff or copying the container upper read-write storage layer (OverlayFS) before termination.
    3. Extract process memory directly from the host operating system via /proc/<PID>/mem for the container isolated process, or use nsenter to execute forensic tools within the container namespaces.
    4. Audit Kubernetes API server logs to trace service account token abuse, unauthorized pod deployments, and cluster privilege escalation.
Loading diagram...
Cloud Snapshot Acquisition & Isolated Forensic Investigation Workflow
Test Your Knowledge

An incident handler must quarantine a running AWS EC2 instance for volatile triage and also cut an already established command-and-control session immediately. Which containment plan accounts for AWS connection tracking?

A

Terminate the EC2 instance immediately through the AWS Management Console

B

Stop the EC2 instance to preserve disk blocks and trigger automatic hypervisor memory dumps

C

Attach a quarantine security group for new flows and apply an approved stateless network ACL or equivalent provider control when the tracked connection must be interrupted immediately

D

Power down the virtual network interface (ENI) using the instance's local operating system command line

Test Your Knowledge

A threat actor compromises an AWS IAM user's credentials and invokes sts:AssumeRole to obtain temporary security credentials for a high-privilege administrative role. The incident handler detects the intrusion, deactivates the user's static access keys, and resets their console password. However, the attacker continues executing administrative actions using the previously issued temporary security tokens. What step should the handler take to revoke permissions from the previously issued role sessions?

A

Delete the IAM role and recreate it with the same name and permissions

B

Reboot all virtual machines in the account to flush the AWS Security Token Service cache

C

Issue a support ticket to AWS Customer Support requesting manual token invalidation

D

Attach an inline IAM policy to the role with an explicit Deny statement matching the aws:TokenIssueTime condition prior to the revocation timestamp

Test Your Knowledge

During an investigation of a compromised Kubernetes worker node, an analyst discovers that an attacker gained root-level control over the underlying Linux host node from inside a compromised application container. Investigation reveals that the container's deployment specification included the volume mount -v /var/run/docker.sock:/var/run/docker.sock. How did this configuration enable the container escape?

A

The volume mount allowed the container to overwrite the host system's /etc/shadow file directly via symbolic links

B

Mounting the host Docker socket allowed the containerized process to issue direct API commands to the host Docker daemon, enabling the creation of privileged host-level containers

C

The Docker socket mount automatically disabled AppArmor and SELinux kernel modules across the physical cluster

D

The volume mount opened an unauthenticated SSH port on the host node listening on the public internet

Sections you finish are checked off in the contents.