Configuring advanced OCR options

All PDF files imported into Continia Document Capture are OCR-processed in accordance with the settings applicable at the time of processing. The default settings apply if you make no adjustments, but it's possible to customize some of the more advanced settings to suit your needs. This article describes the customization process.

To configure OCR settings

To configure how incoming documents are OCR-processed:

  1. Search () for and select Document Categories.
  2. Open the relevant document category. For example, to open the purchase document category, select the PURCHASE line (not the PURCHASE code itself), and then click Edit on the action bar.
  3. On the OCR Processing FastTab, configure the settings as needed. For more information and recommendations, see Details and recommended settings below.

The table below contains recommendations for some of the fields you can customize using the above guide:

FieldDetails and recommendations
Image Resolution

Enter the number of dots per inch (DPI) to be used by Document Capture when storing OCR-processed files as image files. The entered value must be at least 150 DPI – anything below this returns an error.

The higher the entered value, the better the resolution. However, very high values result in correspondingly large image files that take longer to load in the user interface. The recommended value is 300 DPI, which ensures good resolutions and acceptable file sizes.

Image Color Mode

Specify the color mode of the image files that all imported PDF files are converted into. You can choose between the following options:

  • Black & White - image files consisting of only black and white colors.
  • Gray - image files consisting of only grayscale tones.
  • Color - image files consisting of all original colors.

Selecting the Colour option isn't recommended, as this increases the size of the stored image files and consequently slows down image rendering.

Max. number of pages to process per file

Specify how many pages should be OCR-processed for each imported file, enabling you to reduce the import time and optimize the import process. The last three pages of any imported file are always processed, as they typically contain essential information.

Document Capture imposes an overall limit of 500 pages on document import, meaning that no documents longer than 500 pages can be imported into Document Capture – regardless of what value you enter in this field.

OCR Languages

Add all the languages whose character sets should be recognized by Document Capture when OCR-processing incoming documents.

It's recommended that you limit the number of activated languages to the ones generally used in the documents you import (typically only your own native language and, if relevant, English), as enabling too many languages is likely to lower the overall quality of character recognition.

This setting only applies to on-premises OCR installations.

Process PDF files with XML filesEnable the import of XML files embedded in PDFs (such as ZUGFeRD, Factur-X, and XRechnung). For more information, see Enabling the import of PDF files with embedded XML files (ZUGFeRD, XRechnung).
Process Document Type

Certain vendors send invoices in both PDF and XML. By default, Document Capture treats all attachments as individual documents, but this can result in duplicates. Set the desired option:

  • Both - the default behavior, that is, all attachments are processed as individual documents.
  • PDF - processes the PDF and keeps the XML as a drag-and-drop attachment.
  • XML - processes the XML and keeps the PDF as a drag-and-drop attachment.

For information on the remaining customizable fields on the OCR Processing FastTab, which relate to the automatic splitting of documents, see Splitting documents automatically.