6.5 Data Minimization Techniques

Key Takeaways

  • Minimization has three dimensions: fewer attributes, lower precision, and shorter lifetime.

  • On-device inference and client-side coarsening keep raw sensor data, exact locations, and precise timestamps from ever reaching central servers.

  • Each decimal place removed from a coordinate cuts precision about tenfold: three decimals is roughly 110 meters, two is roughly 1.1 kilometers, and one is roughly 11 kilometers.

  • Abstracting data for the use case (an over-21 flag instead of a birth date, a region instead of an address) is the BoK's model minimization technique.

  • Zero-knowledge proofs and selective-disclosure credentials let a service verify a claim, such as age or income range, without receiving the underlying value.

Last updated: October 2026

6.5 Data Minimization Techniques

Quick Answer: Data minimization means collecting and keeping only what a defined purpose needs, at the lowest precision that still works. Engineers enforce it at the collection boundary: reject unneeded fields, run inference on the device and send only results, coarsen timestamps and locations, and, where possible, verify a fact (such as "over 21") without collecting the underlying data at all. The BoK's own example is to "abstract personal data for a specific use case."

The Engineering Principle of Data Minimization

In modern privacy engineering, data minimization is not a discretionary operational guideline; it is a foundational architectural constraint. International privacy frameworks—including GDPR Article 5(1)(c) ("adequate, relevant and limited to what is necessary"), the privacy-by-design controls in ISO/IEC 27701 (limit collection, limit processing, and minimization objectives), the OECD Collection Limitation Principle, and the NIST Privacy Framework disassociated-processing outcomes (for example, CT.DP-P4: system or device configurations permit selective collection or disclosure of data elements)—expect systems to collect and process only the personal data necessary for a defined, legitimate purpose.

Traditional software architectures defaulted to maximalist collection: logging full request payloads, retaining unbounded telemetry, and aggregating customer data in monolithic relational databases under the assumption that historical data might possess future analytical utility. Privacy engineering reverses this default by establishing three distinct technical dimensions of minimization:

  1. Attribute Minimization: Restricting collection schemas to the bare minimum fields required for transaction execution. If a payment service requires only billing zip code and postal verification, full residential street addresses must be rejected at the API boundary.
  2. Granularity Minimization (Fidelity Reduction): Reducing the precision of collected attributes to the lowest level sufficient for the business use case. Examples include converting exact birthdates to boolean age brackets (e.g., is_over_21: true), truncating IP addresses to network subnets, and coarsening GPS coordinates into broad geographic regions.
  3. Temporal Minimization: Discarding transient intermediate representations immediately upon task completion rather than storing them in persistent caches or write-ahead transaction logs.

Client-Side Data Reduction and Schema Trimming

The most effective location to enforce data minimization is at the collection boundary—before sensitive data ever traverses an external network interface to reach cloud infrastructure.

+-------------------------------------------------------------------------+
|                        CLIENT / EDGE DEVICE                             |
|                                                                         |
|  +--------------------+       +--------------------------------------+  |
|  | Raw Sensor Input   | ----> | Edge Processing Engine               |  |
|  | (Camera, Mic, GPS) |       | (TensorFlow Lite / WebAssembly)      |  |
|  +--------------------+       +--------------------------------------+  |
|                                                  |                      |
|                                                  v                      |
|                               +--------------------------------------+  |
|                               | Data Reduction & Coarsening Filter   |  |
|                               | - Truncate GPS coordinates           |  |
|                               | - Bucket timestamps to UTC day       |  |
|                               | - Hash / drop direct identifiers     |  |
|                               +--------------------------------------+  |
|                                                  |                      |
+--------------------------------------------------|----------------------+
                                                   | Minimized Payload Only
                                                   v
+-------------------------------------------------------------------------+
|                         CLOUD INGRESS GATEWAY                           |
|  +-------------------------------------------------------------------+  |
|  | JSON Schema Validator (additionalProperties: false)               |  |
|  +-------------------------------------------------------------------+  |
+-------------------------------------------------------------------------+

1. On-Device Edge Processing

Rather than transmitting raw audio, high-resolution video, or granular biometric samples to centralized backend servers, privacy-preserving architectures execute machine learning inference locally on the client device using runtimes like TensorFlow Lite, ONNX Runtime Mobile, or browser-based WebAssembly (WASM):

  • Smart Cameras and Computer Vision: Instead of streaming continuous uncompressed video to a central server, an on-device convolutional neural network detects specific object classes (e.g., package delivery, human presence) and transmits only a binary event notification and a localized bounding-box crop, immediately dropping the raw video buffer.
  • Voice Interfaces and Speech Processing: Keyword spotting models execute locally in hardware buffers. The microphone remains isolated from network transmission until an explicit wake-word is confirmed, preventing ambient background audio from being transmitted or logged.

2. Client-Side Coarsening and Bucketing

When telemetry or analytical telemetry is necessary, client code applies transformation functions to lower data fidelity prior to transmission:

  • IP Address Truncation: Client network proxies or client-side SDKs strip the host identifier from IP addresses before sending telemetry (zeroing the final octet of an IPv4 address to yield a /24 network, or zeroing at least the 64-bit interface identifier of an IPv6 address to keep only the /64 prefix; more conservative designs keep only the /48 site prefix).
  • Spatial Coarsening (Geohashing): Raw latitude and longitude coordinates with six decimal places provide roughly 0.1-meter precision. Each decimal place removed cuts precision by a factor of ten: three decimals is about 110 meters (a street block), two decimals about 1.1 kilometers (a neighborhood), and one decimal about 11 kilometers (a town). Short geohashes work the same way: a 5-character geohash cell is about 4.9 km × 4.9 km and a 4-character cell is about 39 km × 19.5 km. A weather feature can work with one decimal place or a 4–5 character geohash without ever revealing a home address.
  • Temporal Bucketing: High-precision timestamps (e.g., millisecond epoch values 1728216123456) enable user activity correlation and behavioral fingerprinting across disparate logs. Systems coarsen timestamps by rounding them to the nearest hour or UTC calendar date.

3. Schema Trimming at Ingress Gateways

API gateways serving ingress traffic must enforce strict schema contracts using JSON Schema or Protocol Buffers (Protobuf). Schemas must explicitly declare additionalProperties: false to reject or strip any undeclared fields injected by compromised clients, legacy mobile apps, or third-party web form scrapers:

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "type": "object",
  "properties": {
    "account_id": { "type": "string", "format": "uuid" },
    "postal_code": { "type": "string", "pattern": "^[0-9]{5}$" },
    "tier_selection": { "type": "string", "enum": ["standard", "premium"] }
  },
  "required": ["account_id", "postal_code", "tier_selection"],
  "additionalProperties": false
}

Abstracting Personal Data for a Specific Use Case

The BoK's example of minimization is abstraction: replace detailed data with the coarser fact a use case actually needs. Ask "what decision does this feature make?" and keep only the input to that decision.

Use CaseDetailed Data (Avoid)Abstracted Data (Prefer)
Age-restricted purchaseFull date of birthis_over_21: true from an age check or credential
Regional pricingStreet addressCountry or sales region
Fraud velocity ruleFull transaction historyCount of transactions in the last hour
Weather forecastGPS coordinates to 6 decimalsCity, or coordinates rounded to 1 decimal (about 11 km)
Fitness summaryHeart rate every secondSession average, peak, and duration
Loyalty tierEvery purchase line itemTotal spend band per year

Abstraction works best when the abstracted value is computed before data leaves the source, on the device or at the edge, so the detailed value is never stored centrally. It also simplifies later obligations: there is less to secure, less to return in an access request, and less to delete.

Minimization Across the Life Cycle

Minimization is not a one-time collection decision. It also applies to:

  • Use: give each service only the fields it needs (projection, views, and column-level permissions).
  • Disclosure: send partners tokens or aggregates instead of raw records (Chapter 8).
  • Retention: delete or aggregate data when the purpose ends (Section 6.4).
  • Logging: keep identifiers out of logs, traces, and error reports (Section 6.2).

Zero-Knowledge Proofs and Selective Disclosure

The ultimate realization of data minimization is the elimination of raw data collection entirely through cryptographic verification. In many verification workflows, an organization does not actually need to possess personal data; it merely needs to confirm that a mathematical statement regarding that data is true.

1. Zero-Knowledge Proofs (ZKPs)

A Zero-Knowledge Proof is a cryptographic protocol allowing a prover to demonstrate to a verifier that a statement P(x)=trueP(x) = \text{true} is mathematically valid without revealing the secret witness xx itself. The proof satisfies three properties:

  • Completeness: If the statement is true and both parties are honest, the verifier will be convinced.
  • Soundness: A cheating prover cannot convince the verifier of a false statement except with negligible probability.
  • Zero-Knowledge: The verifier learns nothing other than the fact that the statement is true.

2. Practical Privacy Engineering Applications

  • Zero-Knowledge Range Proofs (e.g., Bulletproofs): Proving that an applicant's annual salary satisfies a loan eligibility predicate (e.g., annual salary between $50,000 and $120,000) or that an individual's age satisfies age≥21\text{age} \ge 21, without disclosing the exact numerical income or birthdate.
  • zk-SNARKs and zk-STARKs: Generating succinct, non-interactive zero-knowledge proofs for complex identity attributes, allowing a user to prove they belong to a whitelist of authorized citizens without disclosing their name or national identification number.
  • Verifiable Credentials with BBS+ Signatures: Modern decentralized identity standards utilize BBS+ multi-message signatures, which support selective disclosure and cryptographic unlinkability. An issuer signs a credential containing ten distinct attributes (e.g., name, date of birth, driver's license number, organ donor status, address). The user can present a cryptographic proof disclosing only the is_over_21 and organ_donor claims, while mathematically blinding all other attributes. Crucially, the proof contains zero persistent identifiers, preventing verifiers from colluding to track the user across transactions.
Test Your Knowledge

An engineering team is designing a smart doorbell camera system that alerts homeowners to delivery package arrivals. To implement data minimization principles, how should the architecture handle video analysis?

A

Store continuous video recordings indefinitely on the doorbell's local flash storage without retention limits, to avoid transmitting data over public networks.

B

Execute on-device neural network inference to detect package delivery events, transmitting only an event notification and cropped thumbnail while discarding continuous ambient video streams.

C

Stream all raw high-definition video directly to an unencrypted centralized cloud datastore for scheduled batch computer vision processing.

D

Transmit complete uncompressed video frames over TLS to an external computer vision provider and rely on contractual non-disclosure agreements for privacy protection.

Test Your Knowledge

A digital identity platform needs to verify whether applicants qualify for low-income utility subsidies (requiring annual income below $35,000) without collecting, transmitting, or storing applicants' actual financial statements or salary figures. Which technical mechanism provides this capability?

A

Client-side AES-256 encryption where the applicant encrypts their exact tax return and transmits the ciphertext to the utility company.

B

A third-party credit bureau integration that transmits the applicant's complete numerical credit history over a mutual TLS connection.

C

Zero-knowledge range proofs that cryptographically demonstrate income satisfies the threshold predicate without disclosing the underlying monetary value.

D

An API gateway transformation rule that masks the first four digits of the applicant's reported income before saving it in a relational database for later audits.

Test Your Knowledge

An online wine shop asks for full date of birth at checkout only to confirm customers are at least 21. Which change best applies the BoK's guidance to abstract personal data for a specific use case?

A

Verify age through an age-check service or credential and store only an over-21 result with the verification date.

B

Collect the date of birth only from customers who buy more than six bottles.

C

Keep the date of birth but delete it after seven years instead of ten.

D

Encrypt the date of birth field in the orders database with a dedicated key that only the compliance team can use.

Sections you finish are checked off in the contents.