pytesseract, ocr extraction
Budget: $750 – $1,500 USD
We have a text extractor that extract text
Areas:
Identify:
We need to enhance the identifier to detect templates faster and better.
For this we need to utilise a big text extraction of all text in the document and then search for key values in that big text. A pretty simple operation.
Dynamic location of text:
For some areas we need to detect a specific string in the document.
From the coordinates of that string we must calculate the bounding box location
From that location, we know the relational location to the other bounding boxes below the detected string and change Y coordinates on all those bounding boxes.
Then you extract the other strings (here you can also utilize Tesseracts coordinates to find them easier and more precise and that is a confidence enhancer of that the correct value is detected))
Your skills:
- python
- pytesseract
- 10 + years of experience, not less.
Areas:
Identify:
We need to enhance the identifier to detect templates faster and better.
For this we need to utilise a big text extraction of all text in the document and then search for key values in that big text. A pretty simple operation.
Dynamic location of text:
For some areas we need to detect a specific string in the document.
From the coordinates of that string we must calculate the bounding box location
From that location, we know the relational location to the other bounding boxes below the detected string and change Y coordinates on all those bounding boxes.
Then you extract the other strings (here you can also utilize Tesseracts coordinates to find them easier and more precise and that is a confidence enhancer of that the correct value is detected))
Your skills:
- python
- pytesseract
- 10 + years of experience, not less.