2.4 Personal Information Sources, Cross-Border Operations & AI Privacy Risk
Key Takeaways
- Inferred and derived data is personal data under both GDPR and CCPA — organizations often overlook it because they did not collect it directly
- GDPR's extraterritorial reach (Article 3(2)) applies to non-EU organizations that offer goods/services to EU data subjects or monitor their behavior — no EU establishment is required
- Post-Schrems II, organizations using Standard Contractual Clauses for EU-to-third-country transfers must complete a transfer impact assessment evaluating whether the destination country's surveillance laws undermine the safeguards
- AI models can memorize and reproduce specific training examples including personal data — differential privacy and output filtering are key mitigations
- GDPR Article 22 grants data subjects the right not to be subject to decisions based solely on automated processing with legal or similarly significant effects — AI-driven hiring or credit decisions require human intervention and appeal mechanisms
Sources, Types, and Uses of Personal Information
The first step in building a privacy framework is understanding what personal information the organization collects, where it comes from, and how it is used. This is not a one-time exercise — it must be maintained as processing activities evolve.
Personal Information Sources
| Source | Typical Data Types | Common Uses |
|---|---|---|
| Customers (direct) | Name, email, address, payment data, purchase history, support interactions | Service delivery, billing, customer support, marketing (with consent) |
| Employees | Identity documents, payroll, benefits, performance reviews, device logs | HR administration, payroll, benefits, workforce management, compliance |
| Vendors / Partners | Business contact info, due diligence data, contract details | Vendor management, procurement, compliance, contract administration |
| Cookies / Telemetry | IP address, device identifiers, browsing behavior, approximate location | Analytics, personalization, security monitoring, advertising |
| Inferred / Derived data | Preference scores, risk scores, behavioral segments, predictions | Marketing segmentation, fraud detection, product recommendations |
| Third-party data | Purchased lists, enriched demographic data, credit scores | Lead generation, risk assessment, personalization |
Inferred data is particularly important — it is often overlooked because the organization did not "collect" it directly, yet it is personal data if it relates to an identifiable person. Under GDPR, inferred data falls within the scope of personal data. Under CCPA, "inferences drawn from personal information" are explicitly included in the definition of personal information, and consumers have the right to know the sources of such inferences.
Cross-Border Operations and Territorial Scope
When an organization operates across borders — or even when it processes data from individuals in other jurisdictions — it encounters overlapping and sometimes conflicting privacy laws.
Extraterritorial Reach of GDPR
GDPR applies to any organization established in the EU/EEA, regardless of where processing occurs. Critically, it also applies to organizations not established in the EU if they:
- Offer goods or services to EU data subjects (Art. 3(2)(a)), or
- Monitor the behavior of EU data subjects (Art. 3(2)(b))
This means a US-only company with a French-language website selling to EU customers may be subject to GDPR. The key factors are intent (targeting EU customers through language, currency, shipping options) and monitoring (tracking EU users' behavior via cookies or analytics).
Data Localization
Some jurisdictions require personal data to be stored or processed within their borders. Russia's data localization law (Federal Law 242-FZ) requires Russian citizens' personal data to be stored in databases located in Russia. China's PIPL imposes localization requirements for certain categories of data and requires security assessments for cross-border transfers. These requirements can conflict with the organization's desire to centralize data infrastructure and must be evaluated during scope definition.
Cross-Border Transfer Mechanisms
When transferring personal data from the EU to a third country (a country not deemed by the EU Commission to provide "adequate" protection), organizations must use a legal transfer mechanism:
| Mechanism | Description | Key Considerations |
|---|---|---|
| Adequacy decision | The EU Commission determines a third country provides adequate protection | Currently covers UK, Japan, South Korea, and EU-US DPF participants. Simplest mechanism — no additional safeguards needed for covered transfers. |
| Standard Contractual Clauses (SCCs) | Pre-approved contract terms binding the importer to GDPR-level protection | Must complete a Transfer Impact Assessment (TIA) to verify the importer can actually honor the clauses in the destination country's legal environment. Most common mechanism. |
| Binding Corporate Rules (BCRs) | Internal rules adopted by multinational groups for intra-group transfers | Requires DPA approval — a lengthy process. Best for large multinationals with frequent, systematic intra-group transfers. |
| Derogations | Specific exceptions: explicit consent, contractual necessity, important public interest | Limited to specific situations; not suitable for regular, systematic transfers. Reliance on consent is fragile because consent must be freely given and withdrawable. |
The Schrems II decision (July 2020) invalidated the EU-US Privacy Shield and established that organizations using SCCs must assess whether the destination country's surveillance laws (e.g., US FISA Section 702, EO 12333) undermine the contractual safeguards. This is the transfer impact assessment. If the assessment reveals inadequate protection, the organization must adopt supplementary measures (e.g., encryption with keys held outside the destination country, pseudonymization) to restore the level of protection.
AI Privacy Risks in the Business Environment
The use of Artificial Intelligence (AI) introduces privacy risks that traditional privacy frameworks were not designed to address. CIPM candidates must understand these risks as part of Domain I, as AI adoption is now a mainstream business activity that the privacy program must govern.
AI Privacy Risk Table
| Risk | Description | Mitigation Strategy |
|---|---|---|
| Training-data exposure | Personal data used to train models may be exposed if the model is shared, deployed externally, or queried by unauthorized parties | Data minimization in training sets; prefer synthetic or anonymized training data; restrict model access |
| Model memorization | Models can memorize and reproduce specific training examples, including personal data such as names, addresses, or medical records | Differential privacy techniques during training; output filtering; post-training evaluation for memorization |
| Automated decision-making (ADM) | AI systems make or support decisions about individuals (credit, hiring, insurance) without meaningful human review | Transparency about ADM use; human-in-the-loop review; opt-out mechanisms where required (GDPR Art. 22) |
| Inference risk | Models infer sensitive attributes (health status, ethnicity, sexual orientation) from non-sensitive inputs, potentially beyond the original purpose | Purpose limitation enforcement; restrict inference to stated purposes; assess inference risks in DPIAs |
| Vendor AI | Third-party AI services (e.g., cloud-based Large Language Models) process personal data in ways the organization may not fully control or understand | Data Processing Agreements; vendor AI risk assessment; contractual limits on data retention and prohibition on using customer data for vendor model training |
| Profiling and segmentation | AI creates detailed profiles that may be used for purposes beyond the original collection, or in ways that cause harm or discrimination | Purpose limitation enforcement; profiling impact assessments; data subject transparency; bias testing |
GDPR Article 22 and Automated Decision-Making
GDPR Article 22 grants data subjects the right not to be subject to a decision based solely on automated processing that produces legal or similarly significant effects. The data subject has the right to:
- Obtain human intervention in the decision
- Express their point of view
- Contest the decision
This is directly relevant when AI is used for credit scoring, job applicant screening, insurance underwriting, or eligibility determinations. The word "solely" is critical — if a human reviews and can override the AI's recommendation, the decision may not be "solely" automated. However, if the human review is rubber-stamp (automatic approval of AI output without meaningful evaluation), regulators may still consider it "solely" automated.
Vendor AI Due Diligence
When engaging an AI vendor, the privacy program must assess:
- Whether the vendor uses customer data to train its own models (many do, by default)
- Data retention policies — how long inputs and outputs are stored
- Sub-processors — whether the vendor's AI is itself hosted by another provider
- Model deployment — whether the model is shared or dedicated
- Security — encryption, access controls, and certifications
The Data Processing Agreement must explicitly address these points. If the vendor uses customer data for training, the organization must either prohibit this in the contract or ensure that the training data is properly anonymized and that the data subjects have been informed.
Worked Scenario: Cross-Border AI Privacy Risk
A European insurance company wants to use a US-based AI vendor to automate claims assessment. The vendor's service processes claimants' personal data (name, policy number, medical records, claim history) in the United States.
Cross-Border Issue
The company must use Standard Contractual Clauses for the transfer (the US is not covered by a blanket adequacy decision, though the EU-US Data Privacy Framework may cover the vendor if it self-certifies). If using SCCs, the company must conduct a transfer impact assessment to evaluate whether US surveillance laws (post-Schrems II) allow the vendor to adequately protect the data. If the assessment reveals gaps, supplementary measures (e.g., end-to-end encryption with keys held in the EU) may be required.
AI Privacy Risk
The AI model may memorize claimants' medical data from training or inference. The company must verify the vendor's data retention policies — specifically, whether the vendor uses customer data to train its own models. If so, the company must either prohibit training use in the contract or ensure the training data is properly anonymized. The risk of model memorization means that even if the vendor promises not to store inputs, the model itself may reproduce personal data in responses to other users.
ADM Risk
If the AI makes automated claim denials without meaningful human review, the company must provide a human appeal process to comply with GDPR Art. 22. A claims assessor must be able to review the AI's decision, consider the claimant's input, and override the AI's recommendation. The company should also conduct a DPIA covering both the cross-border transfer and the automated decision-making aspects.
Vendor Due Diligence
The company must assess the vendor's security posture, privacy practices, and sub-processor chain. The DPA must explicitly prohibit using the insurer's data for the vendor's own model training without explicit, informed consent. The vendor's sub-processors (e.g., the cloud provider hosting the AI) must be disclosed and approved.
Risk Summary
| Risk Category | Specific Risk | Required Action |
|---|---|---|
| Cross-border transfer | EU-to-US transfer of medical data | SCCs + transfer impact assessment + supplementary measures if needed |
| Model memorization | AI may reproduce claimants' medical data | Contractual prohibition on training use; output filtering |
| ADM (Art. 22) | Automated claim denials | Human review process; appeal mechanism; DPIA |
| Vendor AI | Vendor may use data for own model training | DPA with explicit training prohibition; vendor risk assessment |
| Sub-processors | Cloud provider hosting AI may have separate data practices | Sub-processor disclosure and approval requirements |
A US-only company with no EU establishment launches a French-language website that explicitly targets French customers and uses cookies to track their browsing behavior. Under GDPR Article 3, is the company subject to GDPR?
An organization uses a third-party AI service to screen job applicants. The AI generates a suitability score for each candidate, and the hiring manager routinely approves the AI's recommendation without independent evaluation. Which privacy risk is most directly triggered under GDPR?