7.2 De-identification, Pseudonymization, and Tokenization

Key Takeaways

  • Under GDPR Article 4(5) and Recital 26, pseudonymized data remains legally personal data because it can be re-identified using separately stored additional information, whereas true anonymization irreversibly eliminates identifiability and falls outside GDPR scope.

  • ISO/IEC 20889 and NIST SP 800-188 categorize identifiers into direct identifiers (which uniquely identify a subject on their own) and indirect/quasi-identifiers (which enable linkage attacks when combined with external datasets, as proven by Sweeney's voter registry demonstration).

  • Cryptographic pseudonymization must utilize keyed HMAC (such as HMAC-SHA256) with secret peppers secured in FIPS 140-2/3 Level 3 Hardware Security Modules (HSMs) to prevent offline precomputation, dictionary attacks, and rainbow table reversals.

  • Format-preserving encryption (NIST SP 800-38G) encrypts structured values while preserving their length and character set (e.g., card numbers, SSNs); FF1 is the method to use, because NIST's February 2025 second draft revision drops the FF3 family after published attacks.

  • Vault-based tokenization isolates sensitive cleartext in a centralized mapping database, whereas vaultless tokenization uses deterministic FPE keys; both require strict segregation of duties between token consumer systems and de-tokenization endpoints.

Last updated: October 2026

7.2 De-identification, Pseudonymization, and Tokenization

Quick Answer: De-identification is an overarching operational spectrum ranging from pseudonymization to irreversible anonymization. Under GDPR Article 4(5), pseudonymized data remains personal data because the linkage between the pseudonym and the identity can be re-established using separately held additional information. Only true anonymization—where data subjects are irreversibly non-identifiable by any reasonably likely means (GDPR Recital 26)—exempts data from privacy regulation. To secure transformed data against brute-force and linkage attacks, privacy engineering employs keyed HMACs managed via Hardware Security Modules (HSMs), format-preserving encryption (NIST SP 800-38G FF1), and tokenization architectures with strict segregation of duties.


The Legal and Technical Taxonomy: Anonymization, Pseudonymization, and De-identification

A critical failure mode in privacy engineering is the conflation of pseudonymization with anonymization. While both techniques reduce the risk of direct exposure, their legal definitions, cryptographic assumptions, and residual privacy risks differ fundamentally across international frameworks.

+-----------------------------------------------------------------------------+
|                        THE IDENTIFIABILITY CONTINUUM                        |
|                                                                             |
|  [Directly Identifiable] ----> [Pseudonymized / De-Identified] ----> [Anonymized]     |
|  - Real-world PII             - Reversible via separate key       - Irreversible    |
|  - Direct identifiers         - Indirect identifiers remain       - Zero linkability|
|  - Full GDPR scope            - In GDPR scope (Art 4(5))          - Outside GDPR    |
+-----------------------------------------------------------------------------+

1. General Data Protection Regulation (GDPR)

  • Pseudonymization (Article 4(5)): Defined as "the processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person."
  • Legal Status (Recital 26): Pseudonymized data is still personal data. Because the data controller, processor, or a third party holding the key ("additional information") can reverse the transformation or re-identify individuals through correlation, pseudonymized records fall squarely within the scope of the GDPR. However, pseudonymization is recognized as a primary technical safeguard under Article 25 (Data Protection by Design and by Default) and Article 32 (Security of Processing).
  • Anonymization (Recital 26): Personal data rendered anonymously in such a manner that the data subject is not or no longer identifiable. The GDPR explicitly states that the principles of data protection do not apply to anonymous information. To establish true anonymization, a controller must account for "all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly," considering available technology, processing costs, and the state of the art.

2. ISO/IEC 20889:2018 Framework

ISO/IEC 20889 (Privacy enhancing data de-identification terminology and classification of techniques) provides an international taxonomy for transforming datasets. It positions de-identification as the umbrella term for any process that removes the association between a dataset and the data principal. The standard groups techniques into families:

  • Statistical tools: sampling and aggregation.
  • Cryptographic tools: deterministic encryption, order-preserving encryption, format-preserving encryption, homomorphic encryption, and homomorphic secret sharing.
  • Suppression techniques: masking, local suppression, and record suppression.
  • Pseudonymization techniques: replacing identifiers with pseudonyms such as random or keyed values.
  • Generalization techniques: rounding, top and bottom coding, and combining categories.
  • Randomization techniques: noise addition, permutation, and microaggregation.
  • Synthetic data.

Separately, ISO/IEC 20889 describes formal privacy measurement models, such as k-anonymity and differential privacy, used to measure how much re-identification risk remains after these techniques are applied.

3. NIST SP 800-188 Framework

NIST Special Publication 800-188 (De-Identifying Government Datasets) establishes guidelines for US federal systems. NIST models de-identification as a risk mitigation process rather than a binary state. It categorizes attributes into direct identifiers and indirect identifiers, formalizing re-identification risk assessments based on the availability of auxiliary data, attacker motivation, computational feasibility, and potential harm to data subjects.

DimensionDirectly Identifiable DataPseudonymized Data (GDPR Art 4(5))Truly Anonymized Data (GDPR Recital 26)
ReversibilityNot transformedReversible with access to the key, salt, or token vaultIrreversible under all reasonable computational means
Regulatory ScopeFull legal compliance (GDPR, CCPA, HIPAA)Full GDPR applicability; recognized as a security safeguardExempt from GDPR, CCPA, and statutory privacy rules
LinkabilityDirect and explicitLinkable across datasets sharing identical pseudonymsMathematically unlinkable to real-world identities
Primary RiskImmediate compromise upon exfiltrationOffline brute-force, key exfiltration, linkage attacksReconstruction attacks via high-dimensional auxiliary data

Direct vs. Indirect Identifiers (Quasi-Identifiers) and Linkage Attacks

To engineer effective de-identification pipelines, data architects must classify data attributes based on their discriminatory power:

  1. Direct Identifiers: Attributes that uniquely, explicitly, and unambiguously identify a single individual within a population without needing external corroborating data. Examples include Social Security Numbers (SSNs), National Tax Identification Numbers, passport numbers, biometric templates, full names, and personal phone numbers.
  2. Indirect Identifiers (Quasi-Identifiers / QIs): Attributes that do not uniquely isolate an individual in isolation, but when combined with other quasi-identifiers or correlated with external auxiliary datasets, can pinpoint a specific individual. Examples include 5-digit ZIP codes, birth dates, gender, admission and discharge timestamps, marital status, ethnicity, and job titles.
+-----------------------------------------------------------------------------+
|                        ANATOMY OF A LINKAGE ATTACK                          |
|                                                                             |
|  ["De-identified" Medical Dataset]          [Public Voter Registration List]|
|  (Names & SSNs stripped)                    (Purchased public record)       |
|  +-------------------------------+          +------------------------------+|
|  | ZIP   | DOB        | Diagnosis|          | Name       | ZIP   | DOB     ||
|  |-------|------------|----------|          |------------|-------|---------||
|  | 02138 | 1945-07-31 | Angina   | <======> | Wm. Weld   | 02138 | 1945-07-31||
|  | 02138 | 1945-07-31 | Gastritis|  MATCH   |            |       |         ||
|  +-------------------------------+          +------------------------------+|
|                                                                             |
|  Result: 87.1% of US population is uniquely identified by {ZIP, Gender, DOB}|
+-----------------------------------------------------------------------------+

The Latanya Sweeney Linkage Attack Demonstration

In 1997, computer scientist Dr. Latanya Sweeney executed a landmark re-identification experiment that dismantled the assumption that stripping direct identifiers produces anonymous data.

The Massachusetts Group Insurance Commission (GIC) released what it termed "de-identified" health records for state employees to assist medical researchers. The GIC removed direct identifiers (names, addresses, SSNs), but retained quasi-identifiers including 5-digit ZIP code, gender, and full date of birth.

At the time, William Weld was the Governor of Massachusetts. Sweeney knew that Weld resided in Cambridge, Massachusetts. For $20, Sweeney purchased the public voter registration list for the city of Cambridge. The voter database contained every registered voter's legal name, residential address, 5-digit ZIP code, gender, and date of birth.

Sweeney executed a relational SQL join on the shared quasi-identifier tuple: Target=HealthRecords⋈{ZIP,Gender,DOB}VoterRegistry\text{Target} = \text{HealthRecords} \bowtie_{\{\text{ZIP}, \text{Gender}, \text{DOB}\}} \text{VoterRegistry}

Only one individual in Cambridge matched Governor Weld's specific ZIP code (02138), gender (male), and date of birth (July 31, 1945). Sweeney extracted Governor Weld's complete medical history—including surgical procedures, diagnoses, and medication prescriptions—and mailed the records to his office.

Sweeney subsequently analyzed 1990 US Census summary data, proving mathematically that 87.1% of the population of the United States can be uniquely identified by the three-attribute tuple of {5-digit ZIP code, Gender, Date of Birth}. This empirical proof established that removing names and social security numbers provides zero protection against linkage attacks when quasi-identifiers remain unmanaged.


Cryptographic Pseudonymization Mechanics

Implementing cryptographic pseudonymization requires understanding the operational vulnerabilities of hashing algorithms. Naive hashing of identifiers is one of the most common anti-patterns in enterprise software engineering.

1. Deterministic Hashing vs. Randomized Hashing

  • Deterministic Hashing (H(x)H(x)): Applying an unkeyed hash function (e.g., SHA-256(SSN)) always produces the identical digest for a given input. This permits downstream analytical systems to join disparate tables (e.g., matching customer records between marketing and billing) without decrypting the data. However, deterministic hashing is completely vulnerable to frequency analysis, dictionary attacks, and precomputed rainbow tables.
  • Randomized Hashing (H(x∥nonce)H(x \parallel \text{nonce})): Appending an ephemeral nonce or random initialization vector breaks linkability across records. While this maximizes confidentiality, it destroys relational join capability.

2. Salted Hashing and the Low-Entropy Domain Problem

A salt is a random value appended to an input prior to hashing (H(x∥salt)H(x \parallel \text{salt})). In password storage, unique per-record salts defeat precomputed rainbow tables. However, when applied to pseudonymize structured personal identifiers, standard salting fails against offline brute-force attacks because of the low entropy of input spaces:

  • A US Social Security Number has 9 digits (109=1,000,000,00010^9 = 1,000,000,000 possible values).
  • A GPU cluster computing 101110^{11} SHA-256 hashes per second can try every 9-digit value for one known salt in about 10 milliseconds (109/101110^9 / 10^{11} seconds).
  • If an attacker obtains the salted hashes and the salt column, even per-record salts only multiply the work by the number of records, so a database of a million SSNs falls in hours rather than years. The salt is not secret, so it adds little against such a small input space.
+-----------------------------------------------------------------------------+
|                   KEYED HMAC WITH HSM PEPPER ISOLATION                      |
|                                                                             |
|  +--------------------+                                                     |
|  | Cleartext PII      |                                                     |
|  | (SSN: 123-45-6789) |                                                     |
|  +--------------------+                                                     |
|           |                                                                 |
|           v                                                                 |
|  +-----------------------------------------------------------------------+  |
|  | HMAC Engine (HMAC-SHA256)                                             |  |
|  |   Input: PII + Per-Record Salt                                        |  |
|  |   Secret Key: Pepper (256-bit high-entropy cryptographic key)        |  |
|  +-----------------------------------------------------------------------+  |
|           ^                                                                 |
|           | Key retrieved over mTLS / Key never leaves hardware boundary   |
|  +-----------------------------------------------------------------------+  |
|  | FIPS 140-2/3 Level 3 Hardware Security Module (HSM) / Cloud KMS       |  |
|  | (Root of Trust, Strict Access Policies, Audit Logging)                 |  |
|  +-----------------------------------------------------------------------+  |
|           |                                                                 |
|           v Deterministic Surrogate Token                                   |
|  [e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855]         |
+-----------------------------------------------------------------------------+

3. Keyed HMAC and Pepper Management

To prevent offline dictionary attacks, systems must utilize a keyed Hash-based Message Authentication Code (e.g., HMAC-SHA256) parameterized by a secret key known as a pepper: HMAC(K,m)=H((K⊕opad)∥H((K⊕ipad)∥m))\text{HMAC}(K, m) = H\big((K \oplus opad) \parallel H((K \oplus ipad) \parallel m)\big) Where:

  • KK is a cryptographically random 256-bit key (the pepper).
  • mm is the identifier concatenated with a customer-specific or domain-specific context string.
  • ipadipad and opadopad are inner and outer padding constants.

Pepper Isolation and HSM Architecture:

  • The pepper must never reside in application source code, configuration files, environment variables, or database tables.
  • The pepper must be generated and stored inside a FIPS 140-2/3 Level 3 Hardware Security Module (HSM) or dedicated Cloud Key Management Service (AWS KMS, Google Cloud KMS, Azure Key Vault).
  • Applications invoke the HSM cryptographic API over authenticated mutual TLS (mTLS). An attacker who extracts a full database dump cannot mount an offline brute-force attack because they lack the cryptographic pepper locked within the HSM.

Format-Preserving Encryption (FPE: NIST SP 800-38G)

In enterprise systems, replacing cleartext identifiers with traditional cryptographic ciphertexts introduces severe architectural friction. Standard block ciphers like AES-128 or AES-256 operate in modes such as CBC or GCM, outputting 128-bit blocks of raw binary data. Converting this ciphertext to Base64 or hexadecimal strings increases data length and introduces non-numeric characters.

The Schema Constraint Problem

  • Legacy database schemas define credit card numbers as VARCHAR(16) or NUMERIC(16).
  • Downstream billing services, third-party payment gateways, and mainframe COBOL programs enforce strict regular expression validation masks (^[0-9]{16}$) and Luhn algorithm checksums.
  • Storing standard AES ciphertexts requires altering database schemas, expanding column storage limits, rewriting indexing logic, and breaking downstream consumer integrations.
+-----------------------------------------------------------------------------+
|                      THE SCHEMA CONSTRAINT CHALLENGE                        |
|                                                                             |
|  Input Data:  4532-1188-9923-4412 (16 numeric digits)                       |
|                                                                             |
|  Standard AES-256-GCM:                                                      |
|  Output: 3a7f9c2e0b5d8a1f4e9c7b2a6f1d8e0c... (Binary/Base64, 44+ chars)    |
|  -> BREAKS: Schema limits (VARCHAR(16)), validation regex, legacy systems   |
|                                                                             |
|  Format-Preserving Encryption (NIST SP 800-38G FF1):                        |
|  Output: 4921-8842-1054-9381 (Exactly 16 numeric digits)                   |
|  -> PRESERVES: Radix, string length, database types, downstream interfaces  |
+-----------------------------------------------------------------------------+

NIST SP 800-38G: FF1 (and the Retired FF3 Family)

To solve this challenge, NIST published SP 800-38G (2016), which specified two format-preserving encryption (FPE) methods, FF1 and FF3. After researchers published attacks on FF3, a 2019 draft revision proposed a patched FF3-1, and the second public draft of SP 800-38G Rev. 1 (February 3, 2025) removes the FF3 family entirely, raises FF1's minimum domain size, and forbids floating-point arithmetic in implementations. New designs should use FF1.

  • Radix Arithmetic: FPE operates over a user-defined alphabet or finite domain of radix radix∈[2,216]radix \in [2, 2^{16}]. For a standard 16-digit credit card or 9-digit SSN, radix=10radix = 10. For alphanumeric customer codes, radix=36radix = 36 or 6262.
  • Feistel Network Architecture: FF1 builds a Feistel network around AES. The input string of length nn is split into two halves: AA of length ⌊n/2⌋\lfloor n/2 \rfloor and BB of length ⌈n/2⌉\lceil n/2 \rceil. Across 10 rounds (FF3 used 8), the round function applies AES to one half and adds the result to the other half modulo the radix, so the ciphertext keeps the length and character set of the plaintext. Small domains are weak: NIST requires a minimum domain size, so very short fields (for example, a 4-digit PIN) should not be protected with FPE alone.
  • The Tweak (TT): FPE incorporates an unencrypted external parameter called a tweak. The tweak acts like an initialization vector, ensuring that encrypting identical plaintexts under the same master key yields completely distinct ciphertexts when evaluated under different tweaks (e.g., using tenant_id or column_name as the tweak).
  • Structural Preservation and Luhn Checksums: In payment ecosystems, systems can preserve the first 6 digits (the Bank Identification Number or BIN) and the final 4 digits for routing and customer service display, encrypting only the middle 6 digits using FF1. Advanced FPE implementations can calculate and replace the final check digit to ensure the resulting ciphertext satisfies the Luhn algorithm (mod 10), allowing test transactions to traverse legacy validation pipelines without triggering schema errors.

Tokenization Architectures: Vault-Based vs. Vaultless

Tokenization is the process of exchanging sensitive data for a non-sensitive surrogate value (a "token") that has no mathematical or intrinsic relationship to the original plaintext. Tokenization is universally deployed in Payment Card Industry (PCI DSS) architectures to remove core systems from compliance audit scope.

+-----------------------------------------------------------------------------------+
|                         TOKENIZATION ARCHITECTURAL PATTERNS                       |
|                                                                                   |
|  1. VAULT-BASED TOKENIZATION                                                      |
|  +------------+       +-------------------+       +----------------------------+  |
|  | Client App | ----> | Tokenization API  | <---> | Centralized Token Vault    |  |
|  |            |       |                   |       | Table: Token <-> Cleartext |  |
|  +------------+       +-------------------+       +----------------------------+  |
|  [Random CSPRNG tokens; zero math link; database synchronization bottleneck]      |
|                                                                                   |
|  2. VAULTLESS TOKENIZATION                                                        |
|  +------------+       +-------------------+       +----------------------------+  |
|  | Client App | ----> | Tokenization Node | <---> | KMS / HSM (Key Retrieval)  |  |
|  |            |       | (FPE-FF1 Engine)  |       | Zero Database Lookups      |  |
|  +------------+       +-------------------+       +----------------------------+  |
|  [Stateless algorithmic FPE; global scale; key compromise compromises all tokens] |
+-----------------------------------------------------------------------------------+

1. Vault-Based Tokenization

  • Mechanism: When sensitive data is submitted, the token engine generates a random surrogate value using a cryptographically secure pseudo-random number generator (CSPRNG) or a UUID v4 generator. The engine writes an entry into an encrypted database (the Token Vault) mapping Token <-> Cleartext PII.
  • Security Properties: Vault-based tokenization provides the highest theoretical privacy assurance because there is no mathematical relationship between the token and the plaintext. An attacker in possession of billions of tokens cannot mathematically derive any plaintext value, because the mapping is purely relational.
  • Architectural Trade-offs: The centralized vault represents a massive operational bottleneck. Under high-throughput distributed microservice workloads (e.g., hundreds of thousands of transactions per second across global regions), synchronizing, backing up, sharding, and replicating an ACID-compliant token vault introduces severe latency overhead, high availability risks, and single-point-of-failure exposure.

2. Vaultless Tokenization

  • Mechanism: Vaultless tokenization eliminates the central lookup table entirely. Instead of querying a database, the tokenization service applies deterministic Format-Preserving Encryption (AES-FF1) to the plaintext using a master key and a domain-specific tweak.
  • Security Properties: The token is an encrypted ciphertext formatted to resemble the original data. De-tokenization is achieved by executing the inverse FPE decryption function with the corresponding key and tweak.
  • Architectural Trade-offs: Vaultless tokenization offers near-infinite horizontal scalability, sub-millisecond execution times, and zero database storage overhead. However, its security rests entirely on key management. If the master FPE key is exfiltrated, an attacker can decrypt every token generated under that key across the entire enterprise.

Comparison: Vault-Based vs. Vaultless Tokenization

Architectural AttributeVault-Based TokenizationVaultless Tokenization
Token Generation MethodCSPRNG random generation or sequential indexingDeterministic format-preserving encryption (FF1)
Underlying DatastoreRequires high-performance encrypted relational/NoSQL vaultCompletely stateless; zero database required
Mathematical LinkabilityZero mathematical relationshipAlgorithmic relationship governed by cryptographic key
Horizontal ScalabilityBounded by database write locks and multi-region replicationNear-infinite; stateless nodes scale elastically
Failure / Breach ImpactBreach of vault compromises all stored mapping dataCompromise of master key allows global offline de-tokenization
PCI DSS Scope ReductionMaximally removes downstream systems from audit scopeReduces audit scope if cryptographic keys are strictly segregated

Reversibility Controls and Key Segregation

A resilient tokenization architecture enforces strict segregation of duties between token generation and de-tokenization interfaces:

+-------------------------------------------------------------------------+
|                        DE-TOKENIZATION ISOLATION                        |
|                                                                         |
|  [Untrusted Zone / Analytics / Web App]                                 |
|     |                                                                   |
|     v Operates strictly on surrogate tokens                             |
|  +-------------------------------------------------------------------+  |
|  | Analytics Lake / Reporting Services (ZERO DE-TOKENIZE PERMISSION) |  |
|  +-------------------------------------------------------------------+  |
|                                                                         |
|  ==================== BOUNDARY DEFENSE / mTLS ========================  |
|                                                                         |
|  [Restricted PCI Settlement Zone]                                       |
|     |                                                                   |
|     v Authenticated via mTLS + Attested IAM Role                        |
|  +-------------------------------------------------------------------+  |
|  | De-Tokenization Microservice (Audited, Rate-Limited, HSM-Backed) |  |
|  +-------------------------------------------------------------------+  |
|     |                                                                   |
|     v Emits one-time cleartext over TLS directly to Payment Gateway     |
|  [Acquiring Bank / Payment Processor]                                   |
+-------------------------------------------------------------------------+
  1. Zero-De-tokenization Environments: Analytical clusters, business intelligence platforms, machine learning models, and general web applications must never possess permissions to invoke the de-tokenization API. They process, group, and query data exclusively using surrogate tokens.
  2. Isolated De-tokenization Endpoints: The de-tokenization function is isolated within an audited, highly restricted microservice running in a segregated network zone. Calls to the endpoint require mutual TLS (mTLS), hardware-attested identity certificates, and just-in-time administrative authorization.
  3. Ephemeral De-tokenization: When cleartext is required for transaction settlement (e.g., submitting credit card numbers to an acquiring bank), the de-tokenization proxy decrypts the token at the egress gateway, transmits the payload directly over external TLS to the processor, and immediately discards the cleartext buffer from memory without writing it to disk or application logs.
Loading diagram...
Architectural Comparison of Vault-Based and Vaultless Tokenization Systems
Test Your Knowledge

Under GDPR Article 4(5) and Recital 26, which statement correctly describes the legal status of personal data that has undergone cryptographic pseudonymization?

A

It is classified as truly anonymous data and is therefore entirely exempt from the principles of European data protection law.

B

It is permanently converted into non-personal data provided the hashing algorithm uses a random salt of at least 128 bits.

C

It remains personal data subject to the full scope of GDPR because it can be re-attributed to individuals using additional information kept separately.

D

It can be freely shared across commercial data brokers without consent or contractual safeguards because direct identifiers have been removed.

Test Your Knowledge

In a famous 1997 demonstration, computer scientist Latanya Sweeney re-identified the private medical records of Massachusetts Governor William Weld. What technical vulnerability enabled this linkage attack?

A

A zero-day SQL injection vulnerability in the Massachusetts Group Insurance Commission server.

B

ZIP code, sex, and birth date in the health data matched a public voter list.

C

The encryption key used to protect the medical records was generated using a predictable, low-entropy pseudo-random seed.

D

Governor Weld's medical records had been mistakenly released with his Social Security Number and name left unencrypted.

Test Your Knowledge

An enterprise processing 500,000 credit card transactions per second across multi-region active-active cloud clusters needs to implement tokenization while preserving original 16-digit numeric database schemas and avoiding database synchronization bottlenecks. Which architecture best fulfills these operational criteria?

A

Deterministic unsalted SHA-256 hashing storing the first 16 hexadecimal characters of the digest in the payment table.

B

A centralized relational SQL vault mapping table that locks rows during CSPRNG token creation to maintain global ACID consistency.

C

Standard AES-256 in Galois/Counter Mode (GCM), encoding the resulting binary ciphertext as Base64 strings in a widened column.

D

Stateless vaultless tokenization utilizing Format-Preserving Encryption (NIST SP 800-38G FF1) with master keys secured in Hardware Security Modules.

Sections you finish are checked off in the contents.