12.2 Sensitive Data Masking in Telemetry & Messaging
Key Takeaways
CloudWatch Logs data protection detects sensitive data when events are ingested and masks it at every egress point, including the console, Logs Insights, metric filters, and subscription filters; events ingested before the policy stay unmasked.
Data protection policies use managed data identifiers (pattern matching and machine learning, with checksums such as Luhn for card numbers) plus up to 10 custom data identifiers, each a name and a regex of up to 200 characters.
Viewing unmasked log data requires the logs:Unmask IAM permission; users without it, including broad read-only roles, see asterisks in place of each match.
Amazon SNS message data protection (closed to new customers on April 30, 2026) can audit, de-identify (mask or redact), or deny sensitive data on publish (inbound) or delivery (outbound).
Sensitive data audit findings from both CloudWatch Logs and SNS can be routed to Amazon S3 compliance buckets, CloudWatch log groups, or Amazon Data Firehose to maintain verifiable audit trails for PCI DSS, HIPAA, and GDPR compliance.
12.2 Sensitive Data Masking in Telemetry & Messaging
Software applications routinely generate massive volumes of diagnostic logs, operational traces, and asynchronous event messages. During application debugging, error handling, or high-throughput transaction processing, developers and automated systems frequently log raw request payloads, stack traces, and database queries. Without strict guardrails, these telemetry streams inadvertently capture sensitive Personally Identifiable Information (PII)—such as Social Security numbers (SSNs), credit card primary account numbers (PANs), passport numbers, driver's license details, and cryptographic API tokens.
Exposing sensitive data within logging repositories and messaging brokers introduces severe compliance violations under regulatory standards such as PCI DSS Requirement 3 (protecting cardholder data), HIPAA (safeguarding electronic protected health information), and GDPR/CCPA (protecting consumer privacy). To neutralize these threats without burdening application developers with complex sanitization logic, AWS provides native, inline data protection capabilities: Amazon CloudWatch Logs Data Protection Policies and Amazon SNS Message Data Protection.
CloudWatch Logs Data Protection Architecture
Amazon CloudWatch Logs data protection policies audit and mask sensitive data in log groups. Detection happens as each event is ingested; masking is applied at every egress point (console, GetLogEvents, FilterLogEvents, Logs Insights, metric filters, and subscription filters). CloudWatch Logs keeps the original event, which is why a principal with logs:Unmask can still view it. You can attach one policy to a log group or create an account-level policy (PutAccountPolicy) that covers every existing and future log group in the account; when both exist, identifiers from both apply.
Operational Mechanics: Auditing vs. Masking
A CloudWatch Logs data protection policy supports two distinct operational modes within its policy statements:
- Audit (
Audit):- Scans incoming log events and detects sensitive data patterns without modifying the underlying log text.
- Generates structured audit metrics and emits comprehensive finding logs to designated destinations (such as a separate security log group, an Amazon S3 compliance bucket, or an Amazon Data Firehose delivery stream).
- Allows security teams to quantify PII exposure and establish baseline metrics prior to enforcing active redaction.
- De-identify / Mask (
Deidentify):- Replaces each detected value with asterisks whenever the data leaves CloudWatch Logs.
- Example: A log line containing
User SSN: 123-45-6789is shown to a user withoutlogs:Unmaskwith the whole SSN replaced by asterisks; the digits are not partially revealed. - Logs Insights queries, metric filters, and subscription filters also receive masked values, so downstream consumers process sanitized data. Masking applies only to events ingested after the policy exists.
Managed Data Identifiers vs. Custom Data Identifiers
The data protection policy engine detects sensitive data using two mechanisms:
1. Managed Data Identifiers
AWS provides pre-configured Managed Data Identifiers, which combine pattern matching and machine learning (some also require nearby keywords), that detect standard sensitive data formats across multiple geographic jurisdictions and compliance frameworks:
- Financial Identifiers: Credit card numbers (validated using the Luhn checksum algorithm to eliminate false positives), American Express, Diners Club, Discover, JCB, MasterCard, Visa, US bank routing numbers, International Bank Account Numbers (IBAN).
- Personal Identifiers (PII): US Social Security numbers (SSNs, validating format, valid area numbers, and range exclusions), US individual taxpayer identification numbers (ITIN), passport numbers (US, UK, Germany, Canada), driver's license numbers, postal addresses, email addresses, phone numbers.
- National Identifiers: UK National Insurance Numbers (NINO), Canadian Social Insurance Numbers (SIN), French INSEE, German National Identity Cards.
- Security Credentials: AWS Access Key IDs (
AKIA...), AWS Secret Access Keys, SSH private keys, PGP private keys, JSON Web Tokens (JWT).
2. Custom Data Identifiers
Organizations frequently handle proprietary or domain-specific sensitive data that standard managed identifiers cannot classify—such as internal employee badge IDs, patient medical record numbers (MRNs), or custom loyalty card numbers. CloudWatch Logs supports Custom Data Identifiers:
- Name and Regex: Each custom identifier is a name (up to 128 characters) and a regular expression (up to 200 characters) matching the exact format of the proprietary data, for example
EMP-[0-9]{6}-[A-Z]{2}. Omit^and$anchors, because the value usually sits in the middle of a log line. - Limits: Up to 10 custom identifiers per policy, defined in the policy's
Configurationblock and referenced by name. Unlike Amazon Macie, CloudWatch Logs custom identifiers have no keyword or proximity settings, so write specific patterns to limit false positives.
Data Protection Policy JSON Structure
The following policy audits findings to an S3 security bucket and masks the same identifiers, including one custom identifier. The Audit and Deidentify statements must list exactly the same data identifiers:
{
"Name": "payment-service-data-protection",
"Description": "Audit and mask card numbers, SSNs, secret keys, and employee IDs",
"Version": "2021-06-01",
"Configuration": {
"CustomDataIdentifier": [
{ "Name": "EmployeeId", "Regex": "EMP-[0-9]{6}-[A-Z]{2}" }
]
},
"Statement": [
{
"Sid": "AuditSensitiveTelemetry",
"DataIdentifier": [
"arn:aws:dataprotection::aws:data-identifier/CreditCardNumber",
"arn:aws:dataprotection::aws:data-identifier/Ssn-US",
"arn:aws:dataprotection::aws:data-identifier/AwsSecretKey",
"EmployeeId"
],
"Operation": {
"Audit": {
"FindingsDestination": {
"S3": { "Bucket": "centralized-security-compliance-audit-logs" }
}
}
}
},
{
"Sid": "DeidentifySensitiveTelemetry",
"DataIdentifier": [
"arn:aws:dataprotection::aws:data-identifier/CreditCardNumber",
"arn:aws:dataprotection::aws:data-identifier/Ssn-US",
"arn:aws:dataprotection::aws:data-identifier/AwsSecretKey",
"EmployeeId"
],
"Operation": {
"Deidentify": { "MaskConfig": {} }
}
}
]
}
De-anonymization & Unmask Permissions: logs:Unmask
A central design principle of CloudWatch Logs data protection is least-privilege visibility. When data protection is enabled on a log group:
- Standard users, developers, and security analysts possessing standard permissions (
logs:FilterLogEvents,logs:GetLogEvents,logs:StartQuery) only see masked data. - The CloudWatch Logs console displays a prominent banner indicating that data protection is active and that sensitive terms have been redacted.
The logs:Unmask IAM Permission
In authorized operational scenarios—such as a fraud investigation, security incident triage, or regulatory audit—a privileged engineer may require access to the raw, unmasked data. Viewing unmasked data requires the explicit IAM permission logs:Unmask.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowPrivilegedUnmaskWithMFA",
"Effect": "Allow",
"Action": "logs:Unmask",
"Resource": "arn:aws:logs:us-east-1:123456789012:log-group:/aws/production/payment-service:*",
"Condition": {
"Bool": {
"aws:MultiFactorAuthPresent": "true"
}
}
}
]
}
Operational Unmasking Mechanics
- Console Experience: A user with
logs:Unmaskcan click an "Unmask" toggle button within the CloudWatch Logs console. The console prompts the user to confirm temporary unmasking. Once confirmed, the raw data appears in the current session. - API / CLI Experience:
GetLogEvents,FilterLogEvents, andGetLogRecordaccept anunmaskparameter (--unmaskin the CLI). The call succeeds with original values only if the caller haslogs:Unmask; Logs Insights queries use theunmask(@message)function with the same permission. - CloudTrail Auditability: Every invocation of
logs:Unmaskis explicitly logged in AWS CloudTrail. The CloudTrail record captures the exact IAM principal, timestamp, source IP address, target log group, and session context. This non-repudiation mechanism ensures that any attempt to view unmasked PII generates an immutable audit trail for security compliance.
Amazon SNS Message Data Protection
Important
Amazon SNS message data protection is no longer available to new customers (since April 30, 2026). Existing users can keep using it; new designs can inspect payloads in the publisher, or send messages through Lambda or Amazon Comprehend PII detection before publishing.
While CloudWatch Logs protects telemetry, Amazon Simple Notification Service (SNS) provides pub/sub messaging across distributed microservices, serverless architectures, and third-party webhook endpoints. When an upstream payment service publishes transaction notifications, message payloads may contain credit card numbers or customer addresses.
Amazon SNS Message Data Protection scans message payloads in real time as they are published to an SNS topic, evaluating the payload against a topic Data Protection Policy.
Enforcement Actions: Block vs. De-identify
An SNS data protection policy can execute two distinct operational actions upon detecting sensitive data:
- Block (
Deny):- Publisher Rejection: With
DataDirection: Inbound, the publish request is rejected and the publisher receives anAuthorizationError. WithDataDirection: Outbound, delivery to the matching subscriptions is blocked instead. - Zero Ingestion: The message is dropped completely and never delivered to any topic subscribers. This provides an ironclad boundary against sensitive data ingress.
- Publisher Rejection: With
- De-identify / Redact (
Deidentify):- Inline Payload Transformation: The message is accepted by the SNS topic, but the sensitive data strings are automatically masked or redacted before the message is fanned out to downstream subscribers.
- Subscriber Protection: Downstream subscribers (such as Amazon SQS queues, AWS Lambda functions, Amazon Kinesis streams, HTTPS webhooks, or email recipients) receive only sanitized payloads.
SNS Data Protection Policy Example
The following policy configures an Amazon SNS topic to block any message containing Social Security numbers or AWS secret keys from being published, while auditing detections to Amazon CloudWatch Logs:
{
"Name": "payment-topic-data-protection",
"Description": "Block SSNs and AWS secret keys; audit card numbers",
"Version": "2021-06-01",
"Statement": [
{
"Sid": "BlockCredentialsAndSSN",
"DataDirection": "Inbound",
"Principal": ["*"],
"DataIdentifier": [
"arn:aws:dataprotection::aws:data-identifier/Ssn-US",
"arn:aws:dataprotection::aws:data-identifier/AwsSecretKey"
],
"Operation": { "Deny": {} }
},
{
"Sid": "AuditFinancialData",
"DataDirection": "Inbound",
"Principal": ["*"],
"DataIdentifier": [
"arn:aws:dataprotection::aws:data-identifier/CreditCardNumber"
],
"Operation": {
"Audit": {
"SampleRate": "99",
"FindingsDestination": {
"CloudWatchLogs": { "LogGroup": "/aws/vendedlogs/sns/payment-topic-audit" }
}
}
}
}
]
}
Comparison: CloudWatch Logs vs. Amazon SNS Data Protection
| Capability | CloudWatch Logs Data Protection | Amazon SNS Message Data Protection |
|---|---|---|
| Operational Boundary | Telemetry ingestion into log groups. | Real-time message publication to topics. |
| Supported Actions | Audit and Deidentify (Masking). | Audit, Deidentify (Redaction/Masking), and Deny (Block). |
| Blocking / Ingress Denial | No (detects at ingestion and masks on read; never rejects PutLogEvents). | Yes (Rejects sns:Publish call with HTTP 403 error). |
| De-anonymization Mechanism | Yes: Privileged users with logs:Unmask can unmask. | No: Once redacted, downstream subscribers cannot unmask. |
| Managed Identifiers | Full suite (Financial, PII, Credentials, National IDs). | Full suite (Financial, PII, Credentials, National IDs). |
| Custom Regex Identifiers | Supported (name and regex, up to 10 per policy). | Supported (name and regex). |
| Audit Findings Destinations | Amazon S3, CloudWatch Logs, Amazon Data Firehose. | Amazon S3, CloudWatch Logs (/aws/vendedlogs/ log group), Amazon Data Firehose. |
| Availability | Generally available. | Closed to new customers since April 30, 2026. |
Specialty Exam Pitfalls & Architectural Traps
- The Retrospective Masking Fallacy: Believing that enabling a CloudWatch Logs data protection policy will retroactively mask pre-existing log events in the log group. Sensitive data is detected as events are ingested, so events ingested before the policy was set are never masked. To reduce exposure in historical logs, restrict read access to the log group, shorten retention, or delete old log streams after exporting what compliance requires.
- The AdministratorAccess Unmask Assumption: Assuming that an IAM principal with the AWS managed policy
AdministratorAccesscan view unmasked logs by default. WhileAdministratorAccessgrants*on*, if an explicitDenyor Permissions Boundary restrictslogs:Unmask, or if an analyst views the log group without initiating an unmask session, the data remains masked. Always verify whether an analyst possesses explicit, positive authorization forlogs:Unmask. - SNS Block Breaking Critical Pipelines: Configuring an SNS message data protection policy with a
Deny(Block) action on high-priority operational topics. If an application encounters an unexpected error and includes an API key or email address in a critical alert message, the SNS topic will reject the message entirely, dropping the alert and blinding operations teams to an active outage. UseDeidentifyrather thanDenywhen message delivery must be maintained. - Overly Broad Custom Regex: Defining custom data identifier patterns that are too generic (CloudWatch Logs custom identifiers have no keyword or proximity option to narrow them). Unanchored regular expressions frequently trigger massive false positive rates across standard diagnostic telemetry, resulting in unintended masking of harmless hexadecimal identifiers, hash values, and system UUIDs.
A healthcare provider requires all application log groups across 50 AWS accounts to automatically redact Social Security numbers and credit card numbers at log ingestion. Additionally, the Chief Information Security Officer mandates that an immutable audit log detailing every detected occurrence of sensitive data must be delivered to a centralized Amazon S3 bucket in the Security Tooling account for regulatory compliance reporting. Standard DevOps engineers must never view unredacted data. Which architectural solution satisfies these requirements?
In every account, create an account-level CloudWatch Logs data protection policy (deployed with CloudFormation StackSets) whose Audit statement sends findings to the central S3 bucket and whose Deidentify statement masks the same SSN and CreditCardNumber managed data identifiers, and grant logs:Unmask only to compliance officers.
Deploy an Amazon Data Firehose subscription filter on each log group that invokes an AWS Lambda function to execute regex redaction and writes modified logs to Amazon S3.
Enable Amazon Macie across all accounts to scan CloudWatch Logs log groups on a 24-hour schedule and configure Macie finding export to the centralized S3 bucket.
Configure AWS WAF log redaction rules on all Application Load Balancers and attach an IAM policy denying logs:GetLogEvents to DevOps engineers.
A security analyst is investigating a suspected credential compromise. The analyst has been granted the AWS managed policy ReadOnlyAccess along with full AmazonCloudWatchReadOnlyAccess. While searching a payment processing log group in the CloudWatch Logs console, the analyst observes that all credit card numbers appear as asterisks (****************). The analyst requires immediate visibility into the unmasked values to correlate the data with payment gateway logs. What action must the security administrator take?
Attach the AdministratorAccess policy to the analyst's IAM principal and have them re-run the CloudWatch Logs Insights query.
Disable the data protection policy on the log group temporarily so historical log events revert to plaintext.
Grant the analyst explicit IAM permissions for the logs:Unmask action on the specific log group resource ARN and enforce Multi-Factor Authentication.
Instruct the analyst to use the AWS CLI FilterLogEvents command, which automatically bypasses console-based masking.
An e-commerce platform publishes order processing events to an Amazon SNS topic that fans out messages to multiple downstream SQS queues, third-party webhook subscribers, and analytics systems. To comply with PCI DSS, the security team mandates that any message containing credit card primary account numbers (PAN) must be immediately rejected at the publish boundary and prevented from entering the SNS topic or propagating to any subscribers. How should this policy be implemented?
Create an Amazon SQS dead-letter queue on each subscription and deploy a Lambda function to purge messages containing card numbers.
Attach an Amazon SNS Message Data Protection policy to the topic containing a Statement with a Deny operation referencing the CreditCardNumber managed data identifier.
Attach an SNS Topic Policy with a Condition block evaluating StringNotLike on the message body against standard credit card regex patterns.
Configure an Amazon EventBridge rule that intercepts the sns:Publish API call in AWS CloudTrail and invokes an automated remediation Step Function.
A company enables a CloudWatch Logs data protection policy on an existing application log group that contains 180 days of retained diagnostic logs. Following policy activation, a compliance auditor reviews log entries generated 30 days prior and discovers that employee Social Security numbers are completely visible in plaintext. What is the reason for this finding?
The data protection policy requires up to 48 hours to complete a background indexation scan of historical log streams.
The auditor's IAM role inadvertently inherits the logs:Unmask permission from an AWS Organizations Service Control Policy.
Managed data identifiers for Social Security numbers only apply to log streams created by AWS Lambda functions.
CloudWatch Logs data protection policies evaluate and mask log events strictly at ingestion time; pre-existing historical logs remain unmodified.
Sections you finish are checked off in the contents.