12.1 Identifying and Minimizing Privacy Risks in AI, Machine Learning, and Deep Learning

Key Takeaways

  • Models can memorize personal data: researchers extracted verbatim names, phone numbers, and addresses from GPT-2's training data in 2021, so a trained model is not automatically anonymous.

  • Membership inference determines whether a person's record was in the training data, and model inversion reconstructs sensitive attributes from a model's outputs.

  • The EDPB's Opinion 28/2024 says an AI model trained on personal data can be considered anonymous only if personal data cannot be extracted or inferred with reasonable means.

  • Generative AI deployments add prompt leakage, retrieval that ignores document permissions, over-privileged agents, and hallucinated statements about people; OWASP's 2025 LLM Top 10 lists sensitive information disclosure as LLM02.

  • Mitigations include training-data minimization and deduplication, differential privacy in training (DP-SGD), permission-aware retrieval, input and output filtering, no-training contract terms, red-teaming, and governance under NIST AI RMF, ISO/IEC 42001, and the EU AI Act.

Last updated: October 2026

12.1 Identifying and Minimizing Privacy Risks in AI, Machine Learning, and Deep Learning

Quick Summary: Artificial intelligence (AI) systems learn from data, and much of that data is personal. Privacy risk therefore appears in the training data (where it came from and whether it may be used), inside the model (what it memorized), at inference (what attackers can extract and what users type in), and in the outputs (what the model says about people). Section 5.2 covered bias in automated decisions; this section covers the privacy risks of AI as a workplace and product technology.

The BoK lists AI, machine learning (ML), and deep learning under workplace technologies, reflecting how quickly employees and products have adopted them. Deep learning models, including large language models (LLMs), are trained on enormous datasets and contain billions of parameters, which makes their behavior powerful and hard to audit.


Risks Across the AI Life Cycle

Loading diagram...

1. Training Data

  • Lawful basis and purpose limitation: customer data collected to provide a service is often reused to train models without a compatible purpose or consent (Section 7.5).
  • Web scraping: publicly available data remains personal data (Section 6.3). The EDPB's Opinion 28/2024 (December 2024) explains how legitimate interests may or may not support training on such data and lists mitigating measures.
  • Sensitive data: training sets often contain health, financial, and children's data that was never meant for model training.

2. Memorization and Extraction

Neural networks can memorize rare sequences seen during training. In 2021, Carlini and colleagues extracted hundreds of verbatim training examples from GPT-2, including names, phone numbers, email addresses, and physical addresses. Memorization increases with model size and with data that appears many times. A model that can reproduce personal data is itself a store of personal data.

3. Inference Attacks on Models

  • Membership inference: determining whether a specific person's record was in the training set, for example by observing that the model is unusually confident on that record (Shokri and colleagues, 2017). Membership alone can be sensitive if the dataset is, say, patients with a particular disease.
  • Model inversion: reconstructing sensitive features of training data from model outputs (Fredrikson and colleagues showed reconstruction of recognizable faces from a face-recognition model in 2015).
  • Attribute inference: using a model to predict attributes a person never disclosed.

Because of these attacks, the EDPB concluded in Opinion 28/2024 that an AI model trained on personal data can be treated as anonymous only if the likelihood of extracting personal data from it, or obtaining it through queries, is insignificant with reasonable means.

4. Generative AI in Daily Use

  • Prompt leakage: employees paste confidential or personal data into public chatbots. In 2023 Samsung restricted staff use of generative AI after engineers pasted internal source code into a public tool.
  • Retrieval-augmented generation (RAG): a chatbot that searches internal documents can reveal content the user is not permitted to see if retrieval ignores document permissions.
  • Agents with excessive access: AI agents that can read mailboxes or call APIs can exfiltrate data when manipulated by prompt injection hidden in a web page or email.
  • Vector databases: embeddings of documents and conversations are derived personal data that must be secured, inventoried, and deleted on request.
  • Hallucinations about people: models generate confident false statements about real individuals, raising accuracy and defamation concerns (Section 8.2).

The OWASP Top 10 for LLM Applications (2025) lists Prompt Injection as LLM01 and Sensitive Information Disclosure as LLM02, alongside risks such as system prompt leakage, excessive agency, and vector and embedding weaknesses.

5. Rights Requests

Erasure and correction are hard once data is embedded in model weights. "Machine unlearning" research is immature, so practical options include removing the data from future training sets, filtering outputs, retraining on schedule, and avoiding training on data likely to be the subject of requests.


Controls That Minimize AI Privacy Risk

StageControl
DataDocument sources and lawful basis; exclude sensitive and children's data unless justified; scrub direct identifiers; deduplicate training data (reduces memorization)
TrainingDifferentially private training (DP-SGD) bounds what any one record can influence; federated learning keeps raw data on devices
TestingRed-team for training-data extraction, membership inference, and prompt injection before release
DeploymentInput DLP that blocks or redacts personal data in prompts; output filters for personal data; permission-aware retrieval that enforces the user's access rights; least-privilege tools for agents
VendorsEnterprise terms that prohibit training on customer prompts, define retention, and specify data location
GovernanceAI inventory, impact assessments (DPIA and, for certain deployers, the AI Act's fundamental rights impact assessment), model cards, and incident response for AI failures

Governance Frameworks

  • NIST AI Risk Management Framework (AI RMF 1.0, January 2023) with its functions Govern, Map, Measure, and Manage, and the Generative AI Profile (NIST AI 600-1, July 2024), which lists data privacy among the risks unique to or worsened by generative AI.
  • ISO/IEC 42001:2023, a certifiable AI management system standard.
  • EU AI Act, with transparency duties for chatbots and synthetic content from August 2, 2026 and high-risk obligations from December 2, 2027 for Annex III systems (Section 5.2).
Test Your Knowledge

A company fine-tunes a language model on several years of customer support tickets. Testers find that prompting the model with a customer's name sometimes produces that customer's phone number. Which risk does this demonstrate, and which training-side control most directly reduces it?

A

Model drift; retrain the model every week on the newest tickets so that older customer details fade from memory

B

Memorization; scrub identifiers, deduplicate data, or use differentially private training

C

Prompt injection; add a banner warning users not to trust any outputs that mention customer names

D

Algorithmic bias; rebalance tickets across customer demographics

Test Your Knowledge

An internal AI assistant uses retrieval-augmented generation over the company's document store. A junior employee asks about executive compensation and receives figures from a board document she cannot open directly. What is the most appropriate fix?

A

Remove the assistant's citations so that users cannot see which internal documents were used to build each answer.

B

Add a system prompt instructing the model never to discuss compensation or board materials with anyone.

C

Enforce each user's document permissions at retrieval time.

D

Restrict the assistant to senior employees only.

Test Your Knowledge

According to the EDPB's Opinion 28/2024, when can an AI model trained on personal data be considered anonymous?

A

Automatically, because model weights are numbers rather than personal data

B

Whenever the training data was publicly available on the internet

C

Whenever the developer removed names, email addresses, and other direct identifiers from the training set before training began

D

Only when extracting or obtaining personal data from it is insignificantly likely with reasonable means

Sections you finish are checked off in the contents.