You are viewing limited content. For full access, please sign in.

Question

Question

Smart Fields Use Case - Finding page numbers of new documents for splitting

asked on January 29 • Show version history

We have a situation where a form is being scanned physically from paper. When they scan all the papers are stacked together in the scanner so the result is that all documents end up as one file. We want to break the documents back out into the original individual documents.

We tried many prompts without success on getting the page numbers, but I am wondering why AI can not figure this one out. Although it is not the standard use case, the logic feels very simple to me.

For some reason, we always get one result 1, 2, 3 regardless of what we ask for.

There are many ways to determine the first page of a new document and one is that there is this bold text title at the top and also no computer fonts on the pages we want to ignore, just hand writing.

But prompts like this do not work just returning the same result 1, 2, 3. Why 1, 2, 3 and not all the pages I really don't know.

1 0

Replies

replied on January 29

Maybe Quick Fields or Capture Profiles can help accomplish that, if you don't find a Smart Fields prompt that works well?

0 0
replied on February 26

We are trying to move away from Quickfields

0 0
replied on March 3

It looks like you're trying to split larger batches of pages into documents using smart fields (and I'm guessing Workflow?). That part of Quick Fields functionality should stay in Quick Fields for the time being. Any implementation with Workflow and capture profiles or smart fields is going to be less performant. 

0 0
replied on March 3

Quick fields is a lot of overhead and expensive so if your not needing the physical scanning features it would be better to use a server side technology rather than a desktop app.

I feel a language model can do this job easily.

0 0
replied on March 4

Right, the part that's missing to have a full solution is slicing and dicing the batch of pages into individual documents based on the LLM response. We are looking into it, but it's not quite there yet. You can approximate it with Workflow, but depending on the number of pages, it can be somewhat slow and inefficient. 

0 0
replied on March 13 • Show version history

An improvement to Smart Fields was recently released that will improve page detection during extractions.  Please try again, the results should be closer to what you would normally expect.

0 0
replied on May 27 • Show version history

I see some improvement but it is not dependable enough to actually use it yet. For example using the most basic example of asking it to return all page numbers with specific text found on the upper right of the document, it gets about a 30%-50% accuracy finding only a few of the pages with that text.

0 0
You are not allowed to follow up in this post.

Sign in to reply to this post.