7.1 Minimizing Privacy Risk from Aggregation and Inference
Key Takeaways
Aggregation is a Solove information-processing harm: combining data points creates new knowledge about a person that none of the inputs revealed alone.
Inferences drawn from aggregated data are personal data, and the CJEU held in 2022 that data indirectly revealing a special category (such as sexual orientation) is itself special-category data.
A differencing attack subtracts two aggregate answers that differ by one person to recover that person's value, so minimum query sizes must be combined with overlap controls or noise.
Statistical releases commonly suppress small cells (for example, counts of 1 to 10) and check combinations of quasi-identifiers for uniqueness before publication.
Treat every new join between datasets as new processing that needs review, and use separate pseudonyms per domain so datasets cannot be linked casually.
7.1 Minimizing Privacy Risk from Aggregation and Inference
Quick Summary: The BoK asks technologists to "use data analysis and other procedures to minimize privacy risk associated with the aggregation of personal data." Aggregation risk appears in two ways: profiles (joining many data points about one person) and statistics (publishing summaries from which individuals can still be recovered). Both are managed by controlling which datasets may be combined, testing releases for re-identification risk, and limiting or perturbing what queries return.
The Mosaic Effect and Inference
Daniel Solove lists aggregation as an information-processing harm: gathering pieces of information about a person that are innocuous alone but revealing together. Intelligence analysts call this the mosaic effect. Well-known examples show how aggregation turns ordinary data into sensitive knowledge:
| Case | What Was Aggregated | What It Revealed |
|---|---|---|
| Retail purchase modeling (reported 2012) | Purchases of unscented lotion, supplements, and cotton balls | Likely pregnancy and due date, used to send baby-product coupons |
| Strava global heatmap (January 2018) | Millions of individually shared exercise tracks | The layout of military bases and patrol routes in remote areas |
| AOL search logs (2006) | Months of searches under one pseudonymous number | The identity of user 4417749, located by reporters from her searches |
| Credit card metadata study (2015) | Dates and shops of card transactions | Four data points were enough to single out about 90% of people in a dataset of 1.1 million |
Each case involved data that was authorized, pseudonymous, or even voluntarily shared. The harm came from combination and analysis.
Inferences Are Personal Data
Inferences such as "likely pregnant," "probably gay," or "high credit risk" are personal data about the person they describe. The CCPA lists inferences drawn to create a consumer profile as personal information. In the EU, the Court of Justice held in OT (C-184/20, August 2022) that data capable of indirectly revealing a special category, such as a spouse's name that reveals a person's sexual orientation, must be treated as special-category data. A model that infers health conditions from purchases is therefore processing health data, with all the safeguards that requires.
Procedures That Reduce Aggregation Risk in Profiles
- Join governance. Treat every new join between datasets as new processing. Maintain an allowlist of approved joins tied to documented purposes, and require review (often a DPIA update) for new ones. Data contracts can mark fields as "not joinable."
- Domain-specific pseudonyms. Give each domain or dataset its own pseudonym for the same person (for example, an HMAC of the user ID with a different key per domain). Analytics within a domain still works, but datasets cannot be linked without the keys.
- Inference review. Before a model goes live, list what it infers. Block or require explicit consent for inferences about special categories, and do not keep inferred attributes longer than the decision needs.
- Uniqueness testing. Before sharing a dataset, measure how many records are unique on combinations of quasi-identifiers (age, ZIP code, sex, job title). Unique combinations are re-identification risks even with names removed.
- The "motivated intruder" test. Ask whether a reasonably competent person with access to public information and a reason to try could identify someone, as the UK ICO's anonymization guidance recommends.
Procedures That Reduce Aggregation Risk in Statistics
Publishing counts, averages, or dashboards can leak individual values through statistical disclosure.
The Differencing (Tracker) Attack
Suppose an HR dashboard shows total salary by department, and the system refuses queries about fewer than five people.
- Query 1: total salary of the 10 people in Finance = $900,000.
- Query 2: total salary of Finance employees except those hired in March (only Alice was hired in March) = $790,000.
Both queries cover at least five people, yet the difference reveals Alice's salary: $900,000 − $790,000 = $110,000. Minimum query sizes alone do not stop this; systems also need overlap controls (blocking queries whose result sets differ by only a few people), query auditing, or noise (Section 7.4).
Common Disclosure Controls for Published Tables
| Control | How It Works | Example |
|---|---|---|
| Small-cell suppression | Hide counts below a threshold, plus enough other cells to stop back-calculation | Suppressing cells with 1 to 10 people in a health table, then suppressing a second cell in the row so the hidden value cannot be derived from the total |
| Minimum query set size | Refuse queries that match fewer than k records | Dashboards that show nothing for groups under 25 people |
| Rounding and perturbation | Round counts (for example, to the nearest 5) or add small noise | Census-style tables |
| Top and bottom coding | Group extreme values | "Income above $250,000" instead of exact high incomes |
| Differential privacy | Add calibrated noise and track a privacy budget across queries | Usage statistics released with a fixed epsilon per period |
| Access tiers | Give detailed data only to vetted researchers in a secure environment | Research data enclaves with output checking |
Workflow for a Safe Release
- Define the purpose and the minimum level of detail needed.
- Identify direct identifiers and quasi-identifiers; remove or generalize them.
- Run uniqueness and small-cell checks; apply suppression, rounding, or noise.
- Test with a motivated-intruder attempt, including linkage to likely public datasets.
- Document the residual risk and sign-off, and set terms (such as no re-identification attempts) for recipients.
Aggregation risk is the main reason "we removed the names" is never enough. Section 7.3 shows formal models (k-anonymity, l-diversity, t-closeness) for microdata, and Section 7.4 shows how differential privacy bounds what any set of queries can reveal.
An HR analytics tool blocks any query covering fewer than five employees. An analyst queries total bonus for a 12-person team, then the same total excluding employees who joined in June, where only one employee joined in June. What attack does this illustrate, and what control addresses it?
A brute-force attack on analyst passwords, addressed by requiring multifactor authentication on the analytics tool
A differencing attack, addressed by overlap controls, query auditing, or noise added to answers, because minimum query size alone cannot prevent it
A linkage attack on public voter data, addressed by removing employee names from the bonus table before analysis
A homogeneity attack, addressed by raising the minimum query size from five to ten
In 2018, a fitness app's public global heatmap revealed the layout of remote military bases, although each user had chosen to share individual workouts. Which privacy problem best describes this outcome?
Decisional interference, because the heatmap changed which routes users chose for exercise
Breach of confidentiality, because the app broke a promise to keep workouts private
Aggregation, because combining many individually shared data points revealed sensitive information none of them showed alone
Interrogation, because users were pressured by the app's prompts into sharing their workouts publicly
A health department plans to publish monthly case counts by ZIP code, age band, and diagnosis. Several cells contain one or two people. Which step best reduces re-identification risk while keeping the table useful?
Suppress small cells and add complementary suppression or rounding.
Publish the table as is, because aggregate counts by ZIP code contain no names or direct identifiers.
Publish the table only as a PDF instead of a spreadsheet.
Replace each ZIP code with a SHA-256 hash of the ZIP code before publishing the table.
Sections you finish are checked off in the contents.