2.2 Amazon Security Lake & OCSF Log Normalization

Key Takeaways

  • Amazon Security Lake automatically aggregates and normalizes security telemetry across multi-account, multi-region environments into the Open Cybersecurity Schema Framework (OCSF).

  • Native sources (CloudTrail management and S3/Lambda data events, EKS audit logs, Route 53 Resolver query logs, Security Hub CSPM findings, VPC Flow Logs, and AWS WAF logs) are normalized to OCSF and stored as partitioned Apache Parquet.

  • Regional roll-up architecture enables multi-region global organizations to consolidate security telemetry into designated target regions, reducing data transfer costs while maintaining centralized visibility.

  • Security Lake manages AWS Lake Formation permissions and AWS Glue Data Catalog tables, empowering security analysts to execute federated SQL queries in Amazon Athena without data transformation overhead.

  • Third-party SIEM and analytics platforms integrate as subscribers via two architectural patterns: query-based access using Lake Formation or direct data access via SQS event notifications.

Last updated: September 2026

The Enterprise Telemetry Fragmentation Problem

In large-scale enterprise environments spanning dozens of AWS accounts, multiple regions, and hybrid on-premises data centers, security teams confront severe telemetry fragmentation. Each logging system produces records in disparate, proprietary formats:

  • CloudTrail emits nested JSON structures.
  • VPC Flow Logs emit space-delimited text strings.
  • Route 53 Resolver logs emit distinct JSON schemas.
  • GuardDuty findings use the Amazon GuardDuty finding schema.
  • Security Hub findings use the AWS Security Finding Format (ASFF).
  • On-premises firewalls and identity providers generate syslog, CEF, or custom key-value pairs.

Historically, Security Operations Center (SOC) teams had to build, maintain, and scale complex Extract, Transform, Load (ETL) pipelines using AWS Lambda, Amazon Kinesis, and AWS Glue to parse, re-index, and normalize these streams before ingesting them into a Security Information and Event Management (SIEM) tool. These pipelines introduce significant compute costs, processing latency, and operational fragility.

Amazon Security Lake solves this challenge by automatically centralizing security data from multiple AWS accounts and regions into a purpose-built security data lake, normalizing all telemetry into the open-source Open Cybersecurity Schema Framework (OCSF).


The Open Cybersecurity Schema Framework (OCSF)

OCSF is an open-source, vendor-agnostic standard initiated by AWS and leading cybersecurity vendors. It defines a standardized, hierarchical taxonomy for security events, eliminating the need for custom schema parsing.

[OCSF Root]
   └── Categories (e.g., Network Activity, System Activity, Finding)
         └── Event Classes (e.g., DNS Activity, Network Activity, Security Finding)
               └── Standardized Attributes (e.g., src_endpoint, dst_endpoint, actor, severity_id)

Core OCSF Concepts

  • Categories: High-level groupings of event domains (e.g., Category 1: System Activity, Category 2: Findings, Category 3: Identity & Access Management, Category 4: Network Activity).
  • Event Classes: Specific event implementations within a category. For example, Event Class 4001: Network Activity represents connection flows, while Event Class 4003: DNS Activity represents DNS queries and resolutions.
  • Standardized Data Types and Dictionaries: OCSF standardizes IP addresses (ip), port numbers (port), MAC addresses (mac), hostnames (hostname), and status codes. For example, whether a connection was allowed or denied is mapped to standard disposition_id integers (e.g., 1: Allowed, 2: Blocked).
  • Unified Severity: Threat severity across different security products is normalized to a severity_id integer (1 Informational, 2 Low, 3 Medium, 4 High, 5 Critical, 6 Fatal, plus 0 Unknown and 99 Other), enabling consistent correlation across GuardDuty, third-party endpoint detection and response (EDR) agents, and network firewalls.

Amazon Security Lake Core Architecture

Security Lake is deployed across an entire AWS Organization via a designated Delegated Administrator Account (typically the enterprise Security Tooling or Central Security account).

[ AWS Accounts & Regions ]
   ├── CloudTrail Management & Data Events
   ├── VPC Flow Logs
   ├── Route 53 Resolver Query Logs   === (Native Extraction) ===> [ Amazon Security Lake ]
   ├── Amazon GuardDuty Findings                                       │
   └── AWS Security Hub Findings                                       ▼
                                                             [ Normalization Engine ]
                                                                       │
                                                                       ▼
                                                             [ Apache Parquet in S3 ]
                                                             (Partitioned by Source,
                                                              Region, Account, Date)

Automatically Collected Native Sources

When enabled, Security Lake collects its natively supported AWS sources without agents or custom Lambda triggers:

  1. AWS CloudTrail Management Events
  2. AWS CloudTrail Data Events (for Amazon S3 and AWS Lambda)
  3. Amazon VPC Flow Logs
  4. Amazon Route 53 Resolver Query Logs
  5. AWS Security Hub CSPM Findings (which carry GuardDuty, Inspector, Macie, and IAM Access Analyzer findings, converted from ASFF to OCSF)
  6. Amazon EKS Audit Logs
  7. AWS WAFv2 Logs (web ACL traffic logs, which is how edge telemetry reaches an OCSF data lake)

Storage Mechanics: Apache Parquet & Partitioning

Security Lake stores all ingested and normalized data in Amazon S3 using Apache Parquet format. Parquet is an open-source, columnar storage format optimized for analytical queries:

  • Columnar Layout: Queries that examine specific columns (such as src_endpoint.ip or actor.user.name) only read the relevant columns from disk, avoiding costly full-row table scans.
  • Compression: Parquet is a compressed columnar format, so the stored volume is far smaller than the equivalent raw JSON.
  • Deterministic Partitioning: Data is partitioned by source, Region, account ID, and event day, so object prefixes end in a pattern such as region=us-east-1/accountId=111122223333/eventDay=20260929/ (custom sources are written under an ext/ prefix). Partition pruning allows analytical query engines to scan only the exact temporal and structural slices required.

Multi-Region Roll-Up Architecture

Global enterprises operating across multiple AWS regions face architectural trade-offs between centralized visibility and cross-region data transfer costs. Security Lake resolves this via Regional Roll-ups.

Roll-Up Mechanics

  • Contributing Regions: Regions where workloads run and telemetry originates (e.g., eu-west-1, ap-southeast-1, us-west-2).
  • Roll-up Regions: Designated target regions (e.g., us-east-1) that automatically receive replicated security data from contributing regions.

Security Lake uses Amazon S3 cross-region replication under the hood to consolidate telemetry into the roll-up region. This enables security analysts and SIEM systems to execute queries against a single regional endpoint while complying with corporate retention and cross-region egress budgets.

Exam Tip: Data residency compliance laws (such as GDPR in Europe) may prohibit transferring security logs outside their originating jurisdiction. In such scenarios, do NOT configure roll-up replication for regulated regions; leave them as standalone regional data stores and query them locally.


Querying Security Lake via AWS Lake Formation & Amazon Athena

Security Lake registers all normalized Parquet tables in the AWS Glue Data Catalog and enforces data access controls through AWS Lake Formation.

Security analysts can immediately run standard SQL queries in Amazon Athena against OCSF tables without configuring table DDLs or repair scripts:

-- Investigate high-volume blocked outbound traffic across all VPCs
SELECT 
    time_dt,
    src_endpoint.ip AS source_ip,
    src_endpoint.port AS source_port,
    dst_endpoint.ip AS destination_ip,
    dst_endpoint.port AS destination_port,
    traffic.bytes AS bytes_transferred,
    disposition
FROM 
    amazon_security_lake_glue_db_us_east_1.amazon_security_lake_table_us_east_1_vpc_flow_2_0
WHERE 
    disposition = 'Blocked'
    AND traffic.bytes > 10000000
    AND eventDay = '20260929'
ORDER BY 
    traffic.bytes DESC
LIMIT 100;

Because the underlying data is Parquet, Athena scans only the requested columns and matching date partitions, dramatically accelerating query execution while minimizing query costs (charged at $5.00 per TB scanned).


Subscriber Integration Models: Query Access vs. Data Access

Security Lake decouples log collection from security consumption. Downstream consumers (e.g., Splunk, Datadog, Snowflake, OpenSearch, internal SOC tools) register as Subscribers.

Security Lake supports two subscriber consumption models:

Subscriber Access ModelArchitectural FlowPrimary Use CaseSecurity & Authorization Mechanism
Query AccessThe subscriber queries data in place using Amazon Athena or Redshift Spectrum. No data is copied out of S3.Interactive threat hunting, ad-hoc forensics, BI dashboards, Snowflake federated queries.AWS Lake Formation grants fine-grained, column-level and row-level permissions to the subscriber's IAM principal.
Data AccessSecurity Lake sends S3 event notifications via Amazon SQS when new Parquet files are written. The subscriber pulls objects directly from S3.Ingestion into external SIEMs (Splunk Enterprise, Datadog, Sumo Logic) that require local index storage.Security Lake configures an Amazon SQS queue and S3 bucket policy allowing cross-account IAM role assumption with an sts:ExternalId.
[ Security Lake S3 Bucket ]
        │
        ├── (Direct Query via Athena) ─────────> [ Query Subscriber (Snowflake / BI) ]
        │                                          (Lake Formation Permissions)
        │
        ├── New Parquet Object Created
        │        │
        │        ▼
        ├──> [ Amazon SQS Queue ] ─────────────> [ Data Subscriber (Splunk / Datadog) ]
        │        │                                 (Polls SQS, assumes cross-account
        └────────┴── Fetches Parquet File ──────── IAM role with ExternalId) 

Ingesting Custom & Third-Party Log Sources

Enterprise security requires holistic visibility across cloud and on-premises infrastructure. Security Lake supports ingesting Custom Sources (such as Palo Alto firewall logs, Okta authentication logs, or bespoke container application logs):

  1. Define Source: The administrator registers a custom source in Security Lake, specifying the source name and OCSF event class.
  2. S3 Ingestion Location: Security Lake provisions a dedicated S3 bucket prefix and an IAM execution role.
  3. Transformation: The customer's ingestion pipeline (e.g., AWS Lambda, AWS Glue ETL, or Amazon Data Firehose) transforms raw log events into OCSF-compliant Apache Parquet format and writes the objects to the custom source prefix.
  4. Cataloging: Security Lake invokes an AWS Glue Crawler to update partition definitions in the Glue Data Catalog, making custom data instantly queryable alongside native AWS telemetry.
Loading diagram...
Amazon Security Lake Multi-Region Ingestion & Subscriber Architecture
Test Your Knowledge

A chief information security officer (CISO) wants to minimize Amazon Athena query costs when analyzing billions of historical network connection events spanning six months. Which architectural characteristic of Amazon Security Lake directly contributes to lower query scanning costs?

A

It stores records as raw uncompressed JSON files that Athena reads directly via streaming buffers.

B

It converts incoming telemetry into columnar Apache Parquet format with deterministic date and region partitioning, allowing Athena to scan only relevant columns and partitions.

C

It replaces Amazon S3 with Amazon DynamoDB Global Tables configured with on-demand capacity mode.

D

It automatically purges network flow logs older than 14 days into Amazon Glacier Deep Archive where queries are free of charge.

Test Your Knowledge

An enterprise integrates an external third-party SIEM platform with Amazon Security Lake. The SIEM requires notification whenever new VPC Flow Logs are normalized so that it can pull the transformed objects into its indexing cluster. Which Security Lake subscriber pattern should the security architect deploy?

A

Configure an API Gateway WebSocket endpoint that establishes persistent connections to the SIEM cluster.

B

Set up AWS Lake Formation table permissions granting the external SIEM direct IAM full administrator access to all Glue catalog tables.

C

Deploy an AWS Lambda function that reads S3 objects, unzips them to flat CSV files, and sends them via an Amazon SNS email topic.

D

Create a Data Access subscriber in Security Lake, specifying an SQS queue notification workflow and a cross-account IAM role with ExternalId validation for S3 object retrieval.

Test Your Knowledge

An international retail enterprise operates AWS workloads in us-east-1, eu-west-1, and ap-southeast-1. European legal counsel mandates that security logs originating in the EU must not leave the EU jurisdiction due to strict data residency regulations. How should the security engineer configure Amazon Security Lake?

A

Designate us-east-1 as a roll-up region for ap-southeast-1, but do not configure eu-west-1 to replicate to us-east-1, keeping eu-west-1 as an independent, locally queried region.

B

Configure a single roll-up region in us-east-1 and apply an S3 Bucket Policy in eu-west-1 that denies s3:GetObject to non-EU IP ranges.

C

Disable Amazon Security Lake in eu-west-1 and rely exclusively on local CloudWatch log groups with no central retention.

D

Deploy an AWS Transit Gateway inter-region peering connection to encrypt log data before it crosses international borders into us-east-1.

Test Your Knowledge

A security operations team wants to ingest logs from on-premises perimeter firewalls into Amazon Security Lake so analysts can correlate network edge traffic with AWS workloads using Amazon Athena. What is the required implementation sequence?

A

Install the CloudWatch unified agent directly on the physical on-premises firewalls and configure them to stream syslog to /aws/security-lake/default.

B

Export firewall logs to an on-premises SFTP server and configure AWS DataSync to mirror the files to an Amazon EFS file system mounted on Athena.

C

Register a Custom Source in Security Lake specifying the OCSF event class, transform the firewall logs into OCSF-compliant Parquet format, and write the files to the provided Security Lake S3 bucket prefix.

D

Ingest firewall logs into an Amazon RDS MySQL database and use an AWS DMS replication task to stream CDC events directly into Security Lake.

Sections you finish are checked off in the contents.