You are viewing limited content. For full access, please sign in.

Discussion

Discussion

Email Archive not extracting text on PDF attachemnts

posted on May 27

Hello everyone, I'm using Email Archive in my AP workflow. When the emailed attachment is directly from the vendor, EA seems to extract text just fine. If the PDF is a scanned document and then sent through EA, there is no text generated. I can, however, generate text with LF client. I'm looking for a way that I can generate text on these documents without having to use DCC since I need the text in my workflow as soon as the document is uploaded. Is there a way to configure EA to use the same engine as the LF client to ensure all documents get text? 

0 0
replied on May 27

Email Archive currently only supports retrieving native text from attachments. I created an item in our feature backlog to support OCR attachments. 

1 0
replied on May 27

There are two settings for getting text from documents, Retrieve Text and OCR. Retrieve Text will just retrieve text that is embedded in the electronic file. OCR will try and read the pixels on a page and create text. In Import Agent you will see these two settings on the Processing tab of the job. Does yours have both selected?

0 0
replied on May 27

ya, that's what I thought too. Import agent has this option, Email Archive does not. At least my Email Archive does not.

Version 12.0.2510.614

0 0
You are not allowed to follow up in this post.

Sign in to reply to this post.