Skip to main content

Setting up and applying OCR

  • February 3, 2026
  • 1 reply
  • 331 views

Rosie Clarke
Forum|alt.badge.img+6

 

Edition

 

Optical Character Recognition (OCR) is available for Starter Plus, Professional, Professional Plus and Enterprise edition customers in New Generation.

OCR will be available for EoP customers in a future EoP Release.

With the v8.4 release, the OCR Settings page in Classic is visible for reference but can no longer be changed.

Introduction

 

OCR can make the text within images searchable .

When included as part of an ingest or enrichment policy, OCR converts an image of printed text into searchable data in your Preservica archive. As of November 2025, English is the default language for OCR in New Generation.

Note: OCR in Preservica uses the Tesseract optical character recognition engine and is intended to extract the text from high-quality images of type-written material. It is not intended for handwriting recognition.  

An AI disclaimer must be agreed to before the feature can be configured.

At this time, OCR can only be applied on content as it’s ingested, it is not possible to apply OCR on content already in the archive

You can layer both OCR and PII detection to identify digitized images that have PII when they are being ingested.

The extracted text can be accessed via API or through the New Generation explorer screen or search by adding the Metadata → XIP → “Full text (Partial March)” field to the Grid View. It doesn’t display nicely if very long but you can download search results to better see it in a CSV.

 

Permissions

 

You need to have the Manager or Config manager role to set up the OCR enrichment policy.

You need Read Content permissions on any content you are evaluating with the OCR enrichment policy.

Before you start

 

OCR in Preservica works with the following file formats and their associated PUIDs: 

  • Tiff/EXIF: fmt/353, x-fmt/387 

  • JPG2000: x-fmt/392 

  • PDF: fmt/16, fmt/17, fmt/18, fmt/19, fmt/20 

Set up

 

  1. Navigate to the Settings > Enrichment page

Image showing the enrichment page
  1. Select the Edit policy button

  2. Select “On ingest” as your preferred Mode

  3. Acknowledge the AI disclaimer

  4. Create a Profile

    1. Name your Profile

    2. Select the OCR option

      Image showing the Enrichment policy editor - Profiles screen

       

  5. Press the Save button and then the Next button to move to the Mapping page

  6. Complete the Mapping step:

    1. Select the OCR profile from the Profile dropdown list

    2. Select the options for Mapping your Profile:

      1. Select “Everywhere” if you want OCR applied to all newly ingested content within your system

      2. Select a specific Folder location if you want OCR applied only to content ingested in that location

      3. Select a specific security tag or tags if you want OCR applied only to content ingested with that security tag

      4. These options can be used together to narrow how and where your OCR policy is applied. In the example below, the OCR Profile has been mapped to content ingested into the Marketing folder if it has the security tag “open” or “public” - all other content will be ignored by the OCR tool.

Image showing the Enrichment policy editor - Map profile screen

 

c. Press Save and then the Next button to review your enrichment policy

  1. Review your policy set-up and select Apply policy to save your changes

    1. A notification message will pop up confirming that your policy has been saved and will be applied to your archive

Considerations

 

The OCR process using Tesseract is a very resource intensive activity that may affect system performance, so please follow these guidelines when running: 

  • OCR requires high quality images with printed text. Low quality images or handwritten text will produce inferior results or none at all. 

  • At this time, OCR can only be applied on content as it’s ingested, it is not possible to apply OCR on content already in the archive

  • By its nature the OCR process can take a significant time to run, all images need to be converted, evaluated and scanned. Therefore it can take multiple days to run but this will vary depending on the number and quality of source files. 

  • It is not currently possible to search for text within an image that has had OCR applied to it
    We are planning a further enhancement which will support this

 

1 reply

Forum|alt.badge.img+3
  • Known Participant
  • May 14, 2026

If a PDF is OCR’d before being uploaded, do you then need to use credits to re-OCR the document?