7.3 Microdata Protection: k-Anonymity, l-Diversity, and t-Closeness

Key Takeaways

  • k-Anonymity guarantees that every combination of quasi-identifier values in a sanitized microdata table appears in at least k records, partitioning the dataset into equivalence classes of size greater than or equal to k.

  • k-Anonymity is implemented through generalization (substituting granular values with broader categories or intervals) and suppression (withholding outlier records or highly distinguishing attributes).

  • k-Anonymity is vulnerable to homogeneity attacks (when all records in an equivalence class share the same sensitive attribute value) and background knowledge attacks (where external context eliminates candidate values).

  • l-Diversity resolves homogeneity attacks by requiring at least l well-represented sensitive values within each equivalence class (distinct, entropy, or recursive (c, l)-diversity), but remains vulnerable to skewness and similarity attacks.

  • t-Closeness bounds the Earth Mover's Distance (EMD) between an equivalence class's sensitive attribute distribution and the global dataset distribution to no more than threshold t, overcoming semantic similarity and skewness flaws.

Last updated: October 2026

7.3 Microdata Protection: k-Anonymity, l-Diversity, and t-Closeness

Quick Answer: When releasing granular, record-level microdata for research or public analysis, removing direct identifiers is insufficient. Privacy engineering applies statistical disclosure controls: k-Anonymity ensures that every individual's quasi-identifiers match at least k−1k-1 other individuals within an equivalence class. Because k-anonymity fails against homogeneity and background knowledge attacks, l-Diversity mandates that each equivalence class contain at least ll well-represented sensitive attribute values. Because l-diversity fails against skewness and semantic similarity attacks, t-Closeness requires that the probability distribution of sensitive attributes within any equivalence class does not deviate from the overall dataset distribution by more than distance tt using Earth Mover's Distance.


Microdata Structure: Attributes and Equivalence Classes

In data science and privacy engineering, microdata refers to granular datasets containing individual-level records, where each row represents an individual data principal (such as a patient, tax filer, or survey respondent).

+-----------------------------------------------------------------------------+
|                       MICRODATA ATTRIBUTE TAXONOMY                          |
|                                                                             |
|  +--------------------+----------------------------+---------------------+  |
|  | Direct Identifiers | Quasi-Identifiers (QI)     | Sensitive Attributes|  |
|  | (Names, SSNs)      | (Age, Gender, Postal Code) | (Medical Diagnosis) |  |
|  +--------------------+----------------------------+---------------------+  |
|  | [MUST BE REMOVED]  | [GENERALIZED / SUPPRESSED] | [MUST BE PROTECTED] |  |
|  +--------------------+----------------------------+---------------------+  |
+-----------------------------------------------------------------------------+

To apply statistical disclosure controls, the attributes of a microdata table TT are partitioned into four distinct categories:

  1. Direct Identifiers (IDID): Attributes that uniquely single out an individual on their own (e.g., Name, Social Security Number, National Health ID). Direct identifiers must be completely eliminated or cryptographically tokenized prior to release.
  2. Quasi-Identifiers (QIQI): A set of attributes {A1,A2,…,Am}\{A_1, A_2, \dots, A_m\} that, while not unique identifiers individually, can be combined with external auxiliary datasets to re-identify records (e.g., Age, Gender, Postal Code, Occupation, Ethnicity).
  3. Sensitive Attributes (SASA): Attributes containing confidential, private, or stigmatizing information that the data subject does not want associated with their identity (e.g., Medical Diagnoses, Annual Salary, Criminal Records, Political Contributions).
  4. Non-Sensitive Attributes (NSANSA): Ambient operational fields that are neither distinguishing nor sensitive (e.g., database sequence counters, UI display flags).

Equivalence Classes

An equivalence class EE is a subset of records in a table TT that share the exact same values for all quasi-identifier attributes QIQI: E={r∈T∣r[QI]=q}E = \{r \in T \mid r[QI] = q\} Where qq is a specific vector of quasi-identifier values. All records within an equivalence class are indistinguishable from one another based purely on their quasi-identifier attributes.


k-Anonymity: Mathematical Formulation and Implementation

Introduced by Pierangela Samarati and Latanya Sweeney (1998/2001), k-anonymity is the foundational formal model for microdata sanitization.

+-----------------------------------------------------------------------------+
|                       THE k-ANONYMITY PRINCIPLE (k = 3)                     |
|                                                                             |
|  Equivalence Class 1 (Age: [20-29], Gender: Female, ZIP: 941**):            |
|  - Record 1: Age [20-29], Female, 941**, Diagnosis: Hypertension            |
|  - Record 2: Age [20-29], Female, 941**, Diagnosis: Diabetes                |
|  - Record 3: Age [20-29], Female, 941**, Diagnosis: Asthma                  |
|                                                                             |
|  Size of Equivalence Class: |E| = 3 >= k                                    |
|  Maximum Re-identification Probability: 1 / k = 1 / 3 = 33.3%               |
+-----------------------------------------------------------------------------+

Formal Mathematical Definition

A microdata table TT satisfies k-anonymity with respect to a quasi-identifier set QIQI if and only if every equivalence class EE in TT contains at least kk records: ∀r∈T,∣{r′∈T∣r′[QI]=r[QI]}∣≥k\forall r \in T, \quad |\{r' \in T \mid r'[QI] = r[QI]\}| \ge k

Security Guarantee: An adversary who possesses complete knowledge of an individual's quasi-identifiers cannot distinguish that individual from at least k−1k-1 other individuals in the released dataset. The probability of correctly linking a known individual to their specific record is bounded by: P(Re-identification)≤1kP(\text{Re-identification}) \le \frac{1}{k}

Sanitization Mechanisms: Generalization and Suppression

Transforming an identifiable dataset into a k-anonymous dataset requires applying two core transformations:

  1. Generalization: Generalization replaces specific attribute values with broader, less granular values according to predefined Domain Generalization Hierarchies (DGH):
    • Categorical Hierarchies: Generalizing spatial data (e.g., Street Address →\to 5-digit ZIP →\to 3-digit ZIP →\to State →\to Country). 02138⟶0213*⟶021**⟶0****\text{02138} \longrightarrow \text{0213*} \longrightarrow \text{021**} \longrightarrow \text{0****}
    • Continuous / Numerical Bucketing: Aggregating numerical values into intervals (e.g., Age 27 →\to [25-29] →\to [20-39] →\to < 50).
  2. Suppression: Generalizing an entire column to accommodate a handful of unique outliers destroys analytical utility. Suppression withholds specific data points entirely:
    • Record Suppression (Tuple Outlier Dropping): Removing an isolated outlier record (e.g., a single 104-year-old resident in a small postal district) so that the remaining 99.9% of records do not require extreme generalization.
    • Attribute Suppression (Column Dropping): Completely removing a quasi-identifier column if its high cardinality or entropy makes achieving kk-anonymity impossible without destroying the entire table.
    • Cell Suppression: Replacing an individual cell with a null or suppressed indicator (*).

Vulnerabilities of k-Anonymity

While kk-anonymity guarantees syntactic blending of quasi-identifiers, it provides zero mathematical guarantees regarding the diversity of the sensitive attributes (SASA). This gives rise to two critical attack vectors.

1. The Homogeneity Attack

If all kk records within an equivalence class share the same value for a sensitive attribute, an adversary who determines that a victim belongs to that equivalence class discovers the victim's sensitive condition with 100% certainty, regardless of how large kk is.

Equivalence ClassQuasi-Identifiers: AgeQuasi-Identifiers: GenderQuasi-Identifiers: ZIPSensitive Attribute: Medical Diagnosis
Class 1 (Record 1)[30–35]Male9410*Gastric Cancer
Class 1 (Record 2)[30–35]Male9410*Gastric Cancer
Class 1 (Record 3)[30–35]Male9410*Gastric Cancer
Class 1 (Record 4)[30–35]Male9410*Gastric Cancer

Attack Scenario: Bob is known to be a 32-year-old male living in ZIP 94102. An attacker knows Bob visited the hospital. The attacker locates Bob's equivalence class (Class 1, where k=4k=4). Because every record in Class 1 has the diagnosis "Gastric Cancer," the attacker learns that Bob has Gastric Cancer with absolute certainty. The kk-anonymity threshold (k=4k=4) is satisfied, yet privacy is completely compromised.

2. The Background Knowledge Attack

Even if the sensitive attribute values in an equivalence class are not identical, an adversary can leverage external background knowledge to eliminate impossible values, narrowing the sensitive attribute down to the true value.

Attack Scenario: Alice is a 65-year-old female in an equivalence class with k=4k=4 records where the diagnoses are: {Ovarian Cancer, Heart Disease, Diabetes, Testicular Cancer}. An adversary knows that: (1) biologically, Alice cannot have testicular cancer; and (2) Alice is an elite endurance marathon runner whose routine bloodwork publicly shared on social media confirms normal glucose levels and superior cardiac health. By leveraging external knowledge, the adversary eliminates Testicular Cancer, Heart Disease, and Diabetes, concluding that Alice has Ovarian Cancer.


l-Diversity: Mitigating Homogeneity and Its Variants

To defend against homogeneity and background knowledge attacks, Ashwin Machanavajjhala, Daniel Kifer, Johannes Gehrke, and Muthuramakrishnan Venkitasubramaniam (2006) formulated l-diversity.

Core Principle

An equivalence class is said to satisfy l-diversity if it contains at least ll "well-represented" values for each sensitive attribute. A table satisfies ll-diversity if every equivalence class in the table satisfies ll-diversity.

+-----------------------------------------------------------------------------+
|                       THE l-DIVERSITY VARIANTS                              |
|                                                                             |
|  1. DISTINCT l-DIVERSITY:                                                   |
|     At least l distinct values for sensitive attribute SA in class.         |
|     VULNERABILITY: 1 dominant value (98x Flu, 1x HIV, 1x Cancer -> l=3)     |
|                                                                             |
|  2. ENTROPY l-DIVERSITY:                                                    |
|     Entropy H(E) = -sum p(s)*log(p(s)) >= log(l)                            |
|     Requires values to be distributed evenly across the class.               |
|                                                                             |
|  3. RECURSIVE (c, l)-DIVERSITY:                                             |
|     Most frequent value r_1 < c * sum_{i=l}^m r_i                           |
|     Prevents the most frequent value from overwhelming the tail.            |
+-----------------------------------------------------------------------------+

Mathematical Formulations of l-Diversity

  1. Distinct ll-Diversity: Each equivalence class must contain at least ll distinct physical values for the sensitive attribute. Flaw: A class containing 98 records of "Common Cold", 1 record of "HIV", and 1 record of "Tuberculosis" satisfies distinct 3-diversity, yet an adversary knows that any individual in the class has a 98% probability of having the common cold.
  2. Entropy ll-Diversity: Enforces that the empirical entropy of the sensitive attribute distribution within an equivalence class EE is bounded from below: H(E)=−∑s∈Sp(E,s)log⁡p(E,s)≥log⁡(l)H(E) = -\sum_{s \in S} p(E, s) \log p(E, s) \ge \log(l) Where p(E,s)p(E, s) is the fraction of records in EE possessing sensitive value ss. By enforcing entropy ≥log⁡(l)\ge \log(l), the distribution of sensitive values is forced toward a uniform distribution across at least ll categories.
  3. Recursive (c,l)(c, l)-Diversity: Ensures that the most frequent sensitive value does not overwhelm the remaining tail of values. Let r1,r2,…,rmr_1, r_2, \dots, r_m be the frequencies of the sensitive attribute values in EE, sorted in descending order (r1≥r2≥⋯≥rmr_1 \ge r_2 \ge \dots \ge r_m). The class satisfies recursive (c,l)(c, l)-diversity if: r1<c∑i=lmrir_1 < c \sum_{i=l}^{m} r_i Where cc is a user-defined threshold constant. This guarantees that the most frequent value cannot appear more often than cc times the sum of the frequencies of the (m−l+1)(m - l + 1) least frequent values.

Vulnerabilities of l-Diversity: Skewness and Similarity Attacks

Despite enforcing diversity, ll-diversity remains fundamentally vulnerable when attribute semantics and global distributions are considered.

1. The Skewness Attack

When the global distribution of a sensitive condition is heavily skewed, ll-diversity can fail to prevent massive inferential leakage.

Attack Scenario: Consider a medical test for a rare viral disease where the overall prevalence in the general population is 0.1% (1 in 1,000). An equivalence class has 100 individuals, with 50 testing Negative and 50 testing Positive.

  • The class satisfies distinct 2-diversity and high entropy 2-diversity.
  • However, an adversary who determines that Bob belongs to this equivalence class updates their posterior belief of Bob being infected from 0.1% to 50%—a 500-fold increase in risk. Enforcing 2-diversity did not protect Bob from severe privacy harm.

2. The Similarity Attack

ll-Diversity treats sensitive attribute values as distinct, independent categorical symbols. It completely ignores semantic relationships between values. If the distinct sensitive values within an equivalence class all share a common semantic category, privacy is compromised.

Equivalence ClassAgeZIP CodeSensitive Attribute: Medical Diagnosis
Class 1 (Record 1)[40–50]021**Ulcerative Colitis
Class 1 (Record 2)[40–50]021**Crohn's Disease
Class 1 (Record 3)[40–50]021**Diverticulitis

Attack Scenario: Class 1 satisfies distinct 3-diversity because it contains three completely different medical codes. However, all three diagnoses are severe chronic digestive tract inflammatory disorders. An adversary who identifies a victim in Class 1 learns with 100% certainty that the victim suffers from a debilitating gastrointestinal disease, even though the adversary cannot distinguish which specific illness they possess. A similar attack occurs on salary attributes when values are {\$140,000, \$145,000, \$150,000}: all values are distinct, yet all expose that the individual earns a very high executive income.


t-Closeness: Distributional Alignment via Earth Mover's Distance

To overcome both skewness and similarity attacks, Ninghui Li, Tiancheng Li, and Suresh Venkatasubramanian (2007) introduced t-closeness.

+-----------------------------------------------------------------------------+
|                        THE t-CLOSENESS PRINCIPLE                            |
|                                                                             |
|  Global Population Distribution (Q):                                        |
|  [Healthy: 80%] [Flu: 15%] [Cancer: 5%]                                     |
|                                                                             |
|  Equivalence Class Distribution (P):                                        |
|  [Healthy: 78%] [Flu: 16%] [Cancer: 6%]                                     |
|                                                                             |
|  Distance D[P, Q] <= t  (Using Earth Mover's Distance / EMD)                |
|  Adversary gains negligible information beyond global statistical knowledge!|
+-----------------------------------------------------------------------------+

Formal Mathematical Definition

An equivalence class EE is said to satisfy t-closeness if the distance between the marginal probability distribution of the sensitive attribute within the equivalence class (PP) and the marginal probability distribution of the attribute in the entire global dataset (QQ) is no greater than a threshold tt: D[P,Q]≤tD[P, Q] \le t A microdata table TT satisfies tt-closeness if and only if every equivalence class in TT satisfies tt-closeness.

The Metric: Earth Mover's Distance (EMD / Wasserstein Metric)

Standard statistical distances like total variation distance or Kullback-Leibler (KL) divergence treat distances between distinct categorical states as identical. However, to defend against similarity attacks, the distance metric must incorporate semantic distance between values.

Earth Mover's Distance (EMD) formalizes the minimal amount of work required to transform distribution PP into distribution QQ, where work is defined as probability mass multiplied by the ground distance between attribute states: Work(P,Q,F)=∑i=1m∑j=1mfijdij\text{Work}(P, Q, F) = \sum_{i=1}^m \sum_{j=1}^m f_{ij} d_{ij} Where:

  • fijf_{ij} is the probability mass shifted from state ii in PP to state jj in QQ.
  • dijd_{ij} is the ground distance between values ii and jj.

Why EMD Prevents Similarity Attacks:

  • In numerical or ordinal attributes (such as Salary), moving probability mass between $30,000 and $35,000 carries a low ground distance (d=5,000d=5,000), whereas moving mass between $30,000 and $1,000,000 carries an enormous ground distance (d=970,000d=970,000).
  • In hierarchical categorical attributes (such as medical diagnoses mapped to an ICD-10 ontology tree), moving mass between Crohn's Disease and Ulcerative Colitis (sibling nodes under Gastrointestinal Disorders) costs far less work than moving mass between Crohn's Disease and Lung Cancer.
  • By bounding EMD by threshold tt, an equivalence class is forbidden from concentrating mass inside a localized semantic cluster unless the global dataset also exhibits that exact concentration, completely neutralising similarity attacks.

Practical Engineering Trade-offs and Information Loss Metrics

While kk-anonymity, ll-diversity, and tt-closeness provide conceptual guarantees, implementing them in real-world production environments exposes severe practical limitations.

The Curse of Dimensionality (Charu Aggarwal, 2005)

In high-dimensional datasets (e.g., microdata containing 20 or more quasi-identifier attributes such as census demographics, clickstream history, or mobile app usage):

  • As the dimensionality dd increases, data points become exponentially sparse in high-dimensional space.
  • To find kk records that match across 20 quasi-identifiers, algorithms must generalize categories so aggressively (e.g., collapsing all ages to [0-100] and all locations to USA) that almost all analytical utility is destroyed.
  • Aggarwal proved that for high-dimensional data, either kk-anonymity must suppress almost the entire dataset, or the resulting table must be so coarse as to be completely useless for machine learning and statistical modeling.

Information Loss Metrics

Privacy engineers quantify the penalty imposed by generalization and suppression algorithms using formal metrics:

  1. Discernibility Metric (DMDM): Measures the size penalty of equivalence classes, heavily penalizing large classes and completely suppressed tuples: DM=∑E∈Classes∣E∣2+∑r∈Suppressed∣T∣⋅1DM = \sum_{E \in \text{Classes}} |E|^2 + \sum_{r \in \text{Suppressed}} |T| \cdot 1 Minimizing DMDM pushes equivalence class sizes as close to the target threshold kk as possible.
  2. Normalized Certainty Penalty (NCPNCP): Evaluates how much of an attribute's hierarchical domain is covered by a generalized interval. For a numerical or categorical attribute AA generalized to node vv: NCP(v)=∣v∣∣Domain(A)∣\text{NCP}(v) = \frac{|v|}{|\text{Domain}(A)|} The total information loss is the weighted sum of NCP across all attributes and records. A value of 00 indicates raw ungeneralized data, while a value of 1.01.0 indicates total generalization (complete information loss).
Loading diagram...
Progression of Microdata Sanitization Defenses: k-Anonymity to t-Closeness
Test Your Knowledge

A research hospital publishes an inpatient microdata table that satisfies k-anonymity with k = 5. A malicious investigator discovers that their neighbor is a 45-year-old male living in postal code 94110 who was admitted during the study window. Upon inspecting the corresponding equivalence class of 5 records, the investigator observes that all 5 patients were diagnosed with HIV infection. Which vulnerability of k-anonymity does this scenario illustrate?

A

The curse of dimensionality in high-dimensional microdata.

B

A skewness attack caused by global distribution bias.

C

A similarity attack based on a failure of Earth Mover's Distance across semantically related diagnoses.

D

A homogeneity attack, because the class has no diversity in the diagnosis.

Test Your Knowledge

An engineering team sanitizes a clinical trial dataset using distinct 3-diversity. In one equivalence class, the three distinct medical diagnoses for the sensitive attribute are Ulcerative Colitis, Crohn's Disease, and Diverticulitis. Why does this dataset remain vulnerable to privacy compromise despite satisfying distinct l-diversity?

A

Distinct l-diversity requires at least ten unique categorical categories in every equivalence class to prevent brute-force frequency analysis.

B

The dataset was not encrypted using Format-Preserving Encryption prior to applying generalization hierarchies.

C

The distinct sensitive values are all semantically related gastrointestinal inflammatory conditions, allowing an adversary to execute a similarity attack.

D

The equivalence class violates k-anonymity because l-diversity completely supersedes and removes the record grouping constraints.

Test Your Knowledge

In the t-closeness microdata privacy model, why is Earth Mover's Distance (EMD) preferred over total variation distance or Kullback-Leibler (KL) divergence when evaluating the distance between an equivalence class distribution and the global distribution?

A

EMD accounts for the ground distance between related values instead of treating all categories as equally different.

B

EMD is the only mathematical metric that generates differential privacy Laplace noise.

C

EMD operates strictly on binary direct identifiers and ignores quasi-identifier continuous intervals.

D

EMD runs in constant O(1) time complexity, which eliminates the curse of dimensionality for high-dimensional microdata tables.

Sections you finish are checked off in the contents.