Natural Language Regular Expression for UK Bank and SortCode
Budget: £20 – £250 GBP
We are currently using several third party transcription services and are trying to detect entities such as UK credit/debit card numbers, sort-code etc. We use Google DLP however as the transcription is natural laungage, it can contain words like double, triple, dash and/or the verbal representation of the number, word slopping between the digits.
We have been struggling to build a Regex for some time now as a general "catch-all" we also have elastic search at our disposal if needs be.
Taking the uk debit card number for example 1554-1234-1234-1112
Taking the number 1234, it could be 1234, 12 34, 1 2 3 4 or a verbal representation like "one, two, three four", "twelve, thirty four" or "One-thousand, two hundred and thirty four".
We can enable number transcription so we can get 1 double 5 4 dash 1234 hyphen 1234 triple 2
OR
One "double" five four dash one two three four dash one two three four dash "triple" one 2.
We can and do enable number transcription, so in most instances, the second example is less common. With number transcription, there is no guarantee of fixed digits I.e. 1234 may be represented as 1 5 5 4 - 1234 - 1234 - triple 1 2. We are having inconsistent results with this number formatting.
Furthermore, the words themselves can also be slopped with non-sensical data I.e
1 5 5 4 then the next part is 1234 then its 1234 again and the final part is triple 1 2 (This I can't see being picked up well if at all). Typical slop being "Then, final, next, part etc) with between 1-5 words of joining slop (Instead of the hyphen)
The joining chars may also be special chars (hyphens, slashes or whitespace) we may also find commas etc between each number in addition to whitespace 1,5,5,4 1,2,3,4 1,2,3,4 1,1,1,2
To summarise
Were looking for a RegEx that can look for a sequence of numbers in blocks of 4 separated by some form of slop (in single digit, tens, hundreds or thousands), (Or their verbal representations)
Slopped with joining words, whitespace, or special characters
In a relatively unbroken chain of maybe a minimum of 10 whitespace-separated entities and up to maybe a maximum of 20 whitespace-separated entities (16 for the card numbers, four for words like hyphen and/or double)
A non-exhaustive list of examples but should be enough to give a flavor of the problem is below;
1554-1234
1554-1234 one two three four dash triple one 2.
1554-1234-1234-1112
1,5,5,4 1,2,3,4 1,2,3,4 1,1,1,2
1554 1234 1234 1112
15 54 12 34 12 34 11 12
15 54 - 12 34 - 12 34 - 11 12
155 4 12 34 123 4-1112
1 5 5 4 1 2 3 4 1 2 3 4 1 1 1 2
1 554 hyphen 1 2 3 4 - 12 34 1 1 1 2
1 5 5 4 1 2 3 4 1 2 3 4 triple 2
1 double 4 1 2 3 4 1 2 3 4 triple 2
1 double 5 4 dash 1234 hyphen 1234 triple 2
1 5 5 4 then the next part is 1234 then its 1234 again and the final part is triple 1 2
One double five four dash one two three four dash one two three four dash triple one 2.
One thousand five hundred and fifty four then four dash one two three four dash one two three four dash triple one 2.
Now, with the examples above in mind, the problem also exists for uk-sort codes which follow the standard pattern of 00-00-00
It would be great if this could be generalized somehow to pick up various number sequences or other similar generalized ids etc
Such as things like
123/7363-3030-8339-3833
7363-3030-8339-3833/465
While this is not essential, a catch-all would ultimately cut down on re-work.
If this is not possible, 3 separate regular expressions, one for UK sort codes, one for ul card numbers and one for UK account numbers may be supplied
Speed/greediness is not really an issue, we can run this search in isolation, use it in an ES tokenizer or via GoogleDLP. It will typically run on transcriptions of 5000 words / 25000 chars or less
We have been struggling to build a Regex for some time now as a general "catch-all" we also have elastic search at our disposal if needs be.
Taking the uk debit card number for example 1554-1234-1234-1112
Taking the number 1234, it could be 1234, 12 34, 1 2 3 4 or a verbal representation like "one, two, three four", "twelve, thirty four" or "One-thousand, two hundred and thirty four".
We can enable number transcription so we can get 1 double 5 4 dash 1234 hyphen 1234 triple 2
OR
One "double" five four dash one two three four dash one two three four dash "triple" one 2.
We can and do enable number transcription, so in most instances, the second example is less common. With number transcription, there is no guarantee of fixed digits I.e. 1234 may be represented as 1 5 5 4 - 1234 - 1234 - triple 1 2. We are having inconsistent results with this number formatting.
Furthermore, the words themselves can also be slopped with non-sensical data I.e
1 5 5 4 then the next part is 1234 then its 1234 again and the final part is triple 1 2 (This I can't see being picked up well if at all). Typical slop being "Then, final, next, part etc) with between 1-5 words of joining slop (Instead of the hyphen)
The joining chars may also be special chars (hyphens, slashes or whitespace) we may also find commas etc between each number in addition to whitespace 1,5,5,4 1,2,3,4 1,2,3,4 1,1,1,2
To summarise
Were looking for a RegEx that can look for a sequence of numbers in blocks of 4 separated by some form of slop (in single digit, tens, hundreds or thousands), (Or their verbal representations)
Slopped with joining words, whitespace, or special characters
In a relatively unbroken chain of maybe a minimum of 10 whitespace-separated entities and up to maybe a maximum of 20 whitespace-separated entities (16 for the card numbers, four for words like hyphen and/or double)
A non-exhaustive list of examples but should be enough to give a flavor of the problem is below;
1554-1234
1554-1234 one two three four dash triple one 2.
1554-1234-1234-1112
1,5,5,4 1,2,3,4 1,2,3,4 1,1,1,2
1554 1234 1234 1112
15 54 12 34 12 34 11 12
15 54 - 12 34 - 12 34 - 11 12
155 4 12 34 123 4-1112
1 5 5 4 1 2 3 4 1 2 3 4 1 1 1 2
1 554 hyphen 1 2 3 4 - 12 34 1 1 1 2
1 5 5 4 1 2 3 4 1 2 3 4 triple 2
1 double 4 1 2 3 4 1 2 3 4 triple 2
1 double 5 4 dash 1234 hyphen 1234 triple 2
1 5 5 4 then the next part is 1234 then its 1234 again and the final part is triple 1 2
One double five four dash one two three four dash one two three four dash triple one 2.
One thousand five hundred and fifty four then four dash one two three four dash one two three four dash triple one 2.
Now, with the examples above in mind, the problem also exists for uk-sort codes which follow the standard pattern of 00-00-00
It would be great if this could be generalized somehow to pick up various number sequences or other similar generalized ids etc
Such as things like
123/7363-3030-8339-3833
7363-3030-8339-3833/465
While this is not essential, a catch-all would ultimately cut down on re-work.
If this is not possible, 3 separate regular expressions, one for UK sort codes, one for ul card numbers and one for UK account numbers may be supplied
Speed/greediness is not really an issue, we can run this search in isolation, use it in an ES tokenizer or via GoogleDLP. It will typically run on transcriptions of 5000 words / 25000 chars or less
Related categories:
Natural Language
English (UK) Translator
English Spelling
Elasticsearch
Regular Expressions