You are viewing limited content. For full access, please sign in.

Question

Question

"Page 1" is being scanned as "Pagel" due to close spacing

asked on February 25

I'm performing a First Page Identification on an invoice during a classification in Quick Fields and attempting to use an OCR action on the page indexing to identify the first page. The format is "Page 1 of 2" but due to the spacing between the word "Page" and the page number the scan on the first page is reading the number one as a lowercase L and scanning in "Page 1" as the word "Pagel".

 

This is how the indexing appears on the invoice:

The Optimization Style is set to "Accuracy", and we've played around with local enhancements like Smooth using the Grow, Erode, and Sand and Fill settings but haven't been able to achieve a result that will parse the index from the word.

Does anyone have any input or suggestions on how to improve the OCR so we can pull the page number away from the word "Page" and doesn't scan as "Pagel"? 

0 0

Answer

SELECTED ANSWER
replied on February 25 Show version history

There is like no space at all between Page and 1 so it is always going to assume Pagel is a more accurate guess than Page 1. OCR is not intelligent like language models which will also take into account the fact that it says "of 2" and know that it must be "1 of 2".

Assuming the source is a digital document with no skew, what I would do is keep the "Page" text out of the zone, and just capture the X of Y.

Otherwise just look for Pagel of 2 as your identification condition.

0 0

Replies

You are not allowed to reply in this post.
You are not allowed to follow up in this post.

Sign in to reply to this post.