Skip to main content

Setting up and applying OCR

  • February 3, 2026
  • 1 reply
  • 244 views

Rosie Clarke
Forum|alt.badge.img+6

 

Edition

 

Optical Character Recognition (OCR) is available for Starter Plus, Professional, Professional Plus and Enterprise edition customers in New Generation.

OCR will be available for EoP customers in a future EoP Release.

From the v8.4 release, the OCR Settings page in Classic is visible for reference but can no longer be changed.

 

Introduction

 

OCR can make text within digitized images searchable

When included as part of an ingest or enrichment policy, OCR converts an image of printed text into searchable data in your Preservica archive. As of November 2025, English is the default language for OCR in New Gen.

Note: OCR in Preservica uses the Tesseract optical character recognition engine and is intended to extract the text from high-quality images of type-written material. It is not intended for handwriting recognition.  

 

An AI disclaimer must be agreed to before the feature can be configured.

At this time, OCR can only be applied on content as it’s ingested, it is not possible to apply OCR on content already in the archive

You can layer both OCR and PII detection to identify digitized images that have PII when they are being ingested.

The fully indexed text can be displayed via the Grid customization > XIP > Full text field

 

Permissions

 

You need to have the Manager or Config manager role to set up the OCR enrichment policy.

You need Read Content permissions on any content you are evaluating with the OCR enrichment policy.

 

Before you start

 

OCR in Preservica works with the following file formats and their associated PUIDs: 

  • Tiff/EXIF: fmt/353, x-fmt/387 

  • JPG2000: x-fmt/392 

  • PDF: fmt/16, fmt/17, fmt/18, fmt/19, fmt/20 

 

Set up

  1. Navigate to the Settings > Enrichment page

 

Image shows Settings > Enrichment page
  1. Select the Edit policy button

  2. Select “On ingest” as your preferred Mode

  3. Acknowledge the AI disclaimer

  4. Create a Profile

    1. Name your Profile

    2. Select the OCR tool

Images shows selecting the OCR tool
  1. Press the Save button and then the Next button to move to the Mapping page

  2. Complete the Mapping step:

    1. Select the OCR profile from the Format Profile dropdown list

    2. Select the options for Mapping your Profile:

      1. Select “Everywhere” if you want OCR applied to all newly ingested content within your system

      2. Select a specific Folder location if you want OCR applied only to content ingested in that location

      3. Select a specific security tag or tags if you want OCR applied only to content ingested with that security tag

      4. These options can be used together to narrow how and where your OCR policy is applied. In the example below, the OCR Profile has been mapped to content ingested into the Marketing folder if it has the security tag “open” or “public” - all other content will be ignored by the OCR tool.

Images shows the Map Profile settings

c. Press Save and then the Next button to review your enrichment policy

  1. Review your policy set-up and select Apply policy to save your changes

    1. A notification message will pop up confirming that your policy has been saved and will be applied to your archive

 

 

 

Considerations

 

The OCR process using Tesseract is a very resource intensive activity that may affect system performance, so please follow these guidelines when running: 

  • OCR requires high quality images with printed text. Low quality images or handwritten text will produce inferior results or none at all. 

  • At this time, OCR can only be applied on content as it’s ingested, it is not possible to apply OCR on content already in the archive

  • By its nature the OCR process can take a significant time to run, all images need to be converted, evaluated and scanned. Therefore it can take multiple days to run but this will vary depending on the number and quality of source files. 

  • It is not currently possible to search for text within an image that has had OCR applied to it
    We are planning a further enhancement which will support this

 

1 reply

Forum|alt.badge.img+3
  • Known Participant
  • May 14, 2026

If a PDF is OCR’d before being uploaded, do you then need to use credits to re-OCR the document?