6.3 Extracting Personal Data from Publicly Available Sources
Key Takeaways
Publicly accessible personal data remains protected by privacy law, as twelve data protection authorities stressed in their August 2023 joint statement on data scraping.
The robots.txt protocol is an advisory crawler convention; it neither secures data nor provides a legal basis for collecting it.
An organization extracting public data should document a purpose and lawful basis, collect only allowlisted fields, filter sensitive and children's data, pseudonymize early, honor opt-outs, and keep raw scrapes briefly.
When individual notice is impossible, GDPR Article 14(5)(b) still requires measures such as a public notice explaining the collection and how to object.
Defenses against hostile scrapers include rate limiting, TLS fingerprinting (JA3/JA4), behavioral challenges, and limiting what profiles expose to logged-out visitors.
6.3 Extracting Personal Data from Publicly Available Sources
Quick Summary: "It's public" is not a legal basis. Personal data on websites, social networks, public registers, and code repositories remains protected by privacy law, and scraping it at scale creates new risks: profiles that never existed before, re-identification, and use in contexts people never expected. The BoK asks technologists to minimize risk when their own organization extracts public data, and this section also covers how to protect your users from other people's scrapers.
Organizations extract public data for many reasons: search indexing, price monitoring, threat intelligence, academic research, recruiting, and, increasingly, training AI models. Each use raises the same questions: is there a lawful basis, how much personal data is really needed, and what happens to the people whose data is collected?
Automated Web Scraping, Bot Traffic, and Privacy Boundaries
Automated web scraping involves programmatic extraction of content from web properties using headless browsers (e.g., Puppeteer, Playwright), HTTP client libraries, or specialized crawling clusters. While scraping is widely used for search indexing and price monitoring, it poses severe privacy risks when applied to user profiles, community forums, and public directories.
The Technical Role and Limitations of robots.txt
The Robots Exclusion Protocol (robots.txt) is a text file placed at the root of a domain that instructs automated bots which URL paths they should not crawl:
User-agent: *
Disallow: /user/profiles/
Disallow: /api/internal/
Crawl-delay: 10
From a technical and legal standpoint:
robots.txtis an advisory convention, not an access control mechanism. It provides zero technical enforcement: malicious or unauthorized crawlers can simply ignore the file.- Disallowing paths in
robots.txtdoes not make data confidential, nor does allowing paths grant lawful permission or user consent under privacy laws.
Privacy Implications of Public Data Extraction
A common engineering misconception is that data published publicly on the internet is exempt from privacy regulations. Regulatory enforcement bodies (including European DPAs and the FTC) have repeatedly rejected this premise:
- Lack of Lawful Basis: Scraping public social media profiles, facial photos, or forum contributions to build commercial identity databases or train generative AI models violates GDPR Article 6, as data subjects have neither given consent nor reasonably expected mass automated commercial exploitation.
- Re-Identification and Mosaic Attacks: Automated scrapers can aggregate disparate public datasets (e.g., professional resumes, voting registration records, public forum posts, and code repositories). When linked via unique handles or cross-referenced demographic markers, this aggregated profile destroys the contextual privacy the user relied upon when posting in individual spaces.
Technical Perimeter Defenses Against Automated Scraping
Organizations implement multi-layered defenses at the web perimeter to protect user data from automated harvesting:
- TLS Client Fingerprinting (JA3 and JA4): Standard web scraping libraries (e.g., Python Requests, Curl, Go HTTP) produce distinct TLS
Client Hellopacket configurations—including specific cipher suites, extensions, and elliptic curves—that differ markedly from genuine consumer browsers (Chrome, Safari, Firefox). Web Application Firewalls (WAFs) compute the JA3/JA4 cryptographic hash of incoming TLS handshakes, immediately flagging or blocking non-browser signatures. - HTTP/2 and Network Flow Analysis: Inspects HTTP/2 connection settings (such as
SETTINGS_HEADER_TABLE_SIZE,WINDOW_UPDATEsequences, and pseudo-header ordering) to detect headless automation runtimes even when request user-agent headers are spoofed. - Algorithmic Rate Limiting: Deploying Token Bucket or Leaky Bucket algorithms at the API gateway or CDN layer to limit the number of requests per IP address, autonomous system number (ASN), or API token within rolling time windows.
- Interactive Challenge Gates: Triggering cryptographic proof-of-work challenges or CAPTCHA verifications when behavioral heuristics (such as absence of mouse movement, instantaneous form completion, or erratic navigation intervals) suggest automated scraping agents.
Minimizing Risk When Your Organization Extracts Public Data
Data protection authorities have made their expectations clear. In August 2023, twelve members of the Global Privacy Assembly's enforcement working group issued a joint statement on data scraping, reminding social media companies and scrapers that publicly accessible personal data is still subject to privacy law, and followed up in October 2024 with a concluding statement on the safeguards platforms had adopted. In the EU, the EDPB's Opinion 28/2024 on AI models (December 2024) discusses web-scraped training data and lists measures that can tip a legitimate-interests assessment, such as excluding certain sources and honoring opt-outs. A privacy technologist designing a collection pipeline can apply these steps:
| Step | Technique |
|---|---|
| 1. Define the purpose and lawful basis | Write down why the data is needed; in the EU, document a legitimate-interests assessment and, for large-scale or novel scraping, a DPIA. |
| 2. Choose sources deliberately | Exclude sites whose terms or robots.txt prohibit crawling, sites aimed at children, and sources rich in sensitive data (health forums, dating sites) unless the purpose requires them and safeguards exist. |
| 3. Collect only the needed fields | Parse pages into an allowlisted schema (for example, product name and price) and discard names, photos, and contact details at the crawler instead of storing full pages. |
| 4. Filter sensitive and special-category data | Run classifiers to drop or mask health, sexual orientation, religious, political, biometric, and children's data before storage. |
| 5. Pseudonymize or aggregate early | Replace usernames with keyed hashes, or keep only counts and statistics when individual records are not needed. |
| 6. Honor opt-outs | Maintain a suppression list for people and sites that object, and respect machine-readable opt-out signals where they exist. |
| 7. Avoid re-identification by combination | Do not join scraped data with customer records or other datasets unless that combination is part of the documented purpose. |
| 8. Set short retention and provenance | Record where and when each record was collected, and delete raw scrapes once the derived dataset is built. |
| 9. Provide transparency | When individual notice is impossible, GDPR Article 14(5)(b) still requires appropriate measures such as a public notice describing the collection and how to object. |
These steps apply the same minimization logic used everywhere else in the CIPT: reduce identifiability and volume at the earliest point, then govern what remains.
A data intelligence startup uses automated headless browser clusters to scrape public social media profiles, forum discussions, and professional directories to build re-identification graphs. From a privacy engineering and regulatory perspective, why is relying solely on robots.txt compliance insufficient to justify this automated data collection?
The robots.txt file is legally binding under international copyright treaties but automatically grants full privacy exemption if the file is missing from the server root.
The robots.txt protocol is merely an advisory crawling directive for web search indexing, not a legal authorization or privacy waiver for collecting personal data under privacy frameworks.
Privacy regulations like GDPR and CCPA only apply to private databases and expressly permit unrestricted automated harvesting of all publicly accessible web content.
Modern web browsers automatically strip user-identifiable markers from pages whenever robots.txt permits crawling, which eliminates all secondary privacy risks.
A price-comparison company scrapes retailer product pages that sometimes include customer reviews with reviewers' names and photos. The company needs only product names, prices, and star ratings. Which design best minimizes privacy risk?
Extract only product name, price, and rating at the crawler and discard reviewer data before storage.
Keep all scraped data but publish a privacy notice on the company's website describing the scraping program and its sources.
Hash the reviewers' names with SHA-256 and keep the photos for quality checks.
Store the full page HTML so the parser can be improved later, and restrict access to the raw page store to the data engineering team.
A research team scrapes public social media posts about a disease outbreak and cannot contact the authors individually. Under GDPR Article 14, what transparency step is still expected?
Registering the research project and its data sources with the European Data Protection Board before collection.
Taking appropriate measures, including making information about the collection publicly available.
None, because publicly posted data is exempt from the GDPR.
Sending a direct message to every author before any analysis begins, regardless of how much effort or cost it requires.
Sections you finish are checked off in the contents.