4.2 Configure OCR Support for Sensitive Info Types

Key Takeaways

  • Tenant OCR is optional and billed through Microsoft Syntex pay-as-you-go on an Azure subscription; you do not deploy Syntex models, and you do not create separate image SITs—existing SITs, EDM, fingerprints, and trainable classifiers scan extracted text.
  • Configure OCR once at tenant level (Purview Settings > Optical character recognition), then select locations: Exchange, SharePoint, OneDrive, Teams, Windows, and macOS. Settings generally take effect in about an hour.
  • Each standalone image is one transaction; each PDF page is billed separately. Published document-processing OCR pricing is $0.001 per transaction (USD). One scanned image can feed multiple DLP, insider risk, auto-labeling, and records policies at no extra charge.
  • Endpoint OCR is the Devices location of the same tenant setting: Windows and macOS send images to the cloud, default bandwidth is 1,024 MB per device per day, and endpoint/Teams PDF support is image-only—not SharePoint’s hybrid PDF and RAW catalog.
  • Only images uploaded after OCR is enabled are scanned. eDiscovery OCR is a separate case-level Advanced indexing setting and does not use the tenant OCR meter.
Last updated: August 2026

Configure OCR support for sensitive info types

Quick Answer: Tenant optical character recognition (OCR) is an optional, pay-as-you-go switch. After a Global or SharePoint admin attaches Microsoft Syntex billing to an Azure subscription (you do not have to deploy Syntex models), a Compliance admin turns OCR on under Purview Settings > Optical character recognition (OCR), picks locations, and existing SITs start matching text inside images. You do not create “image SITs.”

Without OCR, a scanned passport PDF, a screenshot of a credit card, or a JPEG of a badge is invisible to the DLP condition content contains sensitive information. The SC-401 bullet Configure optical character recognition (OCR) support for sensitive info types wants the billing prerequisite, the location list, endpoint versus service behavior, published size and cache limits, and the eDiscovery OCR trap.

Why SITs need OCR

Sensitive information types, exact data match (EDM) SITs, document fingerprints, and trainable classifiers all operate on text. Image-only PDFs and photos have no extractable text until OCR runs. After you enable OCR at the tenant level, Microsoft Purview applies your existing policies for data loss prevention, records management (auto-apply retention), and insider risk management to image text in the locations you selected. Auto-labeling policies that look for a SIT gain the same extracted text. Microsoft states you do not need separate data classifiers for OCR: once OCR is on, built-in and custom SITs, EDM SITs, fingerprint SITs, and trainable classifiers scan images as well as documents and email.

OCR settings generally take effect about an hour after you turn them on. Plan lab validation around that delay rather than assuming an immediate match on the first test image.

Billing and licensing (published facts only)

OCR is optional because it is metered. Do not invent a hidden E5 checkbox that “includes unlimited OCR.”

StepWhoWhat Microsoft documents
Azure pay-as-you-go subscriptionGlobal adminRequired in the same tenant if one is not already in place
Connect Syntex / document-processing pay-as-you-goGlobal admin or SharePoint admin with owner or contributor rights on the Azure subscriptionMicrosoft 365 admin center pay-as-you-go services. You do not have to set up Microsoft Syntex document-processing models
Configure OCRCompliance administrator, Compliance data administrator, Global administrator, Information Protection, or Information Protection AdminPurview Settings > Optical character recognition (OCR)
Extra Purview license after billingAfter billing information is entered, the Compliance admin configures OCR without extra setup or licensing requirements

Microsoft’s document-processing pay-as-you-go table (USD) lists Optical character recognition at $0.001 per transaction. Purview’s OCR article points to those Syntex billing pages for pricing and defines a transaction as: each standalone image (JPEG, JPG, PNG, BMP, or TIFF) is one scan; each page of a PDF is a separate scan. A 10-page scanned PDF is 10 scans, not one. One scanned image can feed any number of DLP, insider risk, auto-labeling, and records-management policies at no extra charge.

Use the OCR cost estimator (preview) in the same Settings blade before you enable scanning. First estimation can take about 24 hours; the dashboard then refreshes daily. Microsoft documents estimator limits you should not paper over: OCR and the estimator cannot run at the same time; the estimator is usable once, with 90 days to complete a 30-day estimate; dashboard data is deleted 90 days after first enable; embedded-image estimation is Exchange-only; PDF image estimation is Exchange and Teams, not SharePoint/OneDrive (actual SharePoint/OneDrive bills may be higher); Endpoint estimates 4 images per PDF as a published customer average. Download reports before that window expires.

Configure locations and Exchange scope

  1. Sign in to the Purview portal.
  2. Open Settings > Optical character recognition (OCR).
  3. Select the locations where you want to scan images.
  4. Select groups to include or exclude.
  5. Select Done.

Supported locations and which Purview solutions consume the extracted text:

LocationSolutions that act on OCR results
ExchangeDLP; Information protection auto-labeling; Records management auto-apply retention (keywords and SITs)
SharePoint sitesDLP; Insider risk management (SITs and trainable classifiers in images for risk scoring); Records management auto-apply
OneDrive accountsDLP; Records management auto-apply
Teams chat and channel messagesDLP; Insider risk management
Devices (Windows and macOS)DLP; Insider risk management

Default Exchange scope is All sender groups (incoming, internal, and outgoing). To skip inbound mail, switch to Specific sender groups and name internal groups. The Advanced Setting (Only Exchange) checkbox restricts OCR to mail sent outside the organization so neither inbound nor internal mail is OCRed.

DLP policy tips are not supported for images in Exchange. A user who attaches a screenshot of a card number may still be blocked or audited by policy, but do not expect an Exchange policy tip on that image path.

Service OCR versus endpoint OCR

There is no separate “endpoint OCR SKU” in the tenant OCR article. You check Devices in the same tenant OCR settings. Behavior still differs, and SC-401 will test the difference.

Service-side (Exchange, SharePoint, OneDrive, Teams): Microsoft 365 workloads scan images in the cloud. SharePoint and OneDrive support a wide image catalog (BMP, PNG, JPEG, GIF, HEIC, many RAW camera types, scanned and hybrid PDFs, plus embedded images in DOCX/PPTX/XLSX). Exchange supports JPEG/JPG/PNG/BMP/TIFF and scanned PDFs, plus embedded images in DOCX/PPTX/XLSX and RAR/TAR/ZIP/7z with a limit of 20 embedded images per file, including hybrid PDFs. Teams uses the narrower set also used on endpoints: JPEG, JPG, PNG, BMP, TIFF, and PDF (image only).

Endpoint (Windows and macOS): After OCR is on for devices, endpoints send images to the cloud for scanning. File types are JPEG, JPG, PNG, BMP, TIFF, and PDF (image only) — not SharePoint’s RAW/HEIC catalog and not hybrid searchable PDFs. Microsoft publishes:

  • Default bandwidth 1,024 MB of data per device per day for advanced classification scanning. OCR stops on that device when the daily cap is hit unless you raise it in Endpoint DLP settings (or select unlimited bandwidth, which still does not lift the per-image size cap).
  • Network path must allow blob.core.windows.net (wildcard).
  • If you exclude a path in Endpoint DLP settings, OCR does not scan images in those folders.
  • Endpoint DLP settings also state a 50 MB limit on image files when OCR is enabled, which matches the SharePoint/OneDrive/endpoint file-size row in the OCR article.

Do not tell the exam that endpoint OCR is fully offline or that it uses a different SIT catalog. The same tenant classifiers run; the device is another location with stricter file types and a bandwidth governor.

Published image limits and caching

RequirementPublished limit
File size (Exchange, Teams)20 MB maximum
File size (SharePoint, OneDrive, Windows, macOS)50 MB maximum
Resolution50 × 50 px minimum, 16,000 × 16,000 px maximum
Text extractedFirst 2 million characters
Exchange embedded images20 per file
LanguagesMore than 150

Additional published behaviors:

  • Only images uploaded after OCR is enabled are scanned. Turning OCR on does not backfill the historical image library. Validate with a new upload.
  • Caching to reduce cost: small images such as logos and signatures in Exchange are scanned and billed once per unique image across the tenant for a moving five-day window. Endpoint cache is 30 days, local to the device, storing classifier hits and image hash — not customer image bytes. No cache for standalone images in SharePoint and OneDrive; for embedded Office images, if only text is updated, images are not scanned again. Cache keys include image stream hash and size; any mismatch triggers a new OCR.

Adobe’s Microsoft Purview Information Protection support notes apply when you use DLP features with PDF files in Acrobat; if a PDF path fails, check that support article rather than inventing a Purview file-type exception.

eDiscovery OCR is a different switch

Microsoft Purview eDiscovery can OCR images when they are advanced indexed into a review set, using a case-level Search & analytics setting. That OCR:

  • Does not make images searchable during the initial eDiscovery search.
  • Is distinct from tenant OCR and does not incur the tenant OCR meter.

If the exam question is “make DLP match SSNs in screenshots in OneDrive,” the answer is tenant OCR locations, not an eDiscovery case setting. If the question is “reviewers need text from TIFF evidence in a review set,” the answer is case OCR. Communications compliance documents its own OCR path for that solution’s policies; do not confuse it with the Information Protection SIT setting either.

Turn-on checklist for SC-401

  1. Confirm Azure pay-as-you-go plus Syntex billing (Global or SharePoint admin).
  2. Optionally run the cost estimator; download reports before the 90-day window expires.
  3. In Purview Settings, enable OCR and select only the locations you are willing to pay for.
  4. For Exchange, decide inbound versus outbound group scope.
  5. For devices, confirm Endpoint DLP onboarding, bandwidth, path exclusions, and blob.core.windows.net.
  6. Do not clone SITs for images. Validate with a new upload because historical images are not scanned.
  7. Remember the hour delay, the 20 MB / 50 MB caps, the 20 embedded-image Exchange limit, and that a 10-page scanned PDF is 10 transactions.
Loading diagram...
Tenant OCR for SITs versus separate eDiscovery case OCR
Published OCR maximum file sizes (megabytes)
Test Your Knowledge

A Compliance administrator opens Purview Settings to enable OCR so the Credit Card SIT will match numbers inside scanned PDFs in SharePoint. Billing has not been configured. What does Microsoft require before that admin can turn OCR on?

A
B
C
D
Test Your Knowledge

Your estimator shows a 10-page scanned PDF that will be OCR-processed in Exchange. How does Microsoft count that file for OCR transactions?

A
B
C
D
Test Your Knowledge

You enable OCR for Windows devices so Endpoint DLP can detect Social Security numbers in screenshots. Which statement matches Microsoft’s endpoint OCR documentation?

A
B
C
D