You are viewing limited content. For full access, please sign in.

Question

Question

Pattern Matching in Workflow - Different results with Test than Production

asked on August 4 • Show version history

There is a test feature for pattern matching in the Workflow designer. Often this differs from the results you get when the workflow actually runs. Why does this pattern differ?

[^\r\n]+

Test Result, 1 token

Production result: 2 Tokens

It breaks on the space character after "much". There is no line return or new line here and I verified this with Notepad++

0 0

Replies

replied on August 4

Hey Chad,

 

Where are you getting the token from that you're running the Pattern Match on? For instance Capture Profiles will capture multi-line text and put it into a multi-line token, and that can cause some weird behavior if you're not expecting it.

-Kevin

0 0
replied on August 5 • Show version history

This is just the retreive text activity from a CSV file. We are splitting on line return so we can go through the rows, in the same way we go through rows of data returned from a database. The odd thing is that the test works, but the production does not properly split on line return, it splits on spaces at random places in the data.

0 0
replied on August 5

Can you clarify what you mean by the production result being 2 tokens, and if possible, share more of the workflow design so we can see what tokens are going in and out of the activity?

0 0
replied on August 5

When extracting text from CSV files on import, a line break may be added if the line is too long. A page is added every 60 pages as well. 

1 0
replied on August 5

Like word wrap in notepad? How many characters is too long? Can we not insert characters when extracting text?

The screenshots in the OP show the single token output from the test and the 2 tokens output from the production run.

In the workflow monitor tokens are separated by a line return.

0 0
replied on August 5 • Show version history

Is Laserfiche making more modification to the CSV file we don't know about. I swear I am missing space characters now.

I updated their original pattern matching to get the rows from [^\r\n]+ to [0-9]{6}[\s\S]*?(?=\r?\n[0-9]{6}|\z)

This solved the problem of losing the line, but the newly added line returns were still damaging the data.

So I remove the line returns with the regex [^\r\n]+ all values combined. However I am finding missing space character right where the line return was.

The first regex does not remove any space characters and neither does the second. Who removed the space?

Edit: I did realize I can dump all values into a multi-value token and then use a function to add a space between them again. It still concerns me that something other than a human is taking characters out and putting characters into critical CSV files though.

0 0
replied on August 6 • Show version history

The CSV file is not modified on import. But Retrieve Text activity does not use the CSV file, it uses the text pages generated from it. On import to the repository, the CSV file is converted into text pages by inserting line breaks for lines that are too long and page breaks every 60 lines. 

Why not import the CSV as a lookup table instead and use queries?

 

1 0
replied on August 6

That is a good idea. Now that workflow can replace lookup table data. Retrieve document text was the only way to do this for a very long time.

0 0
You are not allowed to follow up in this post.

Sign in to reply to this post.