ARE YOU A PYTHON PRODIGY? If YES, this project is for you. -- 2
Budget: €750 – €1,500 EUR
I am looking for a skilled Python engineer to create an engine which is capable of finding pre-defined Data Points (DP’s) within the PDF file.
Objectives for the system:
1. It has to be completely offline (no online solutions will be accepted);
2. System has to be integrated with LLM (Llama or similar);
3. LLM must be capable on answering questions (via chatbot interface on GUI) on:
a. DP’s; and
b. Text it read while it was scanning the PDF and looking for DP’s;
c. Previously read documents related to the same client.
4. DP’s are to be saved to SQL table;
5. We need a dashboard, to train the engine on finding DP’s in the document.
Current problems:
A. Large part of documents are scanned, but regardless if the document is scanned or not, it’s all in digital form format. It just so happens that some forms were populated digitally, printed, and scanned;
B. As such, scanned documents are:
a. Tilted;
b. There is noise;
c. Zoom levels are different;
C. There are checkboxes;
D. DP’s can:
a. Change location within the document, but it will always have at least one anchor in fixed location relative to DP.
b. Few DP’s will have multiple entries that are true. In other words we will need unique ID per PDF per DP.
c. Some DP’s can be stretched over two pages.
THIS IS A LONG-TERM PROJECT. Project as defined above is to allow me to achieve MVP. Whoever creates me working engine, will become partner in business. In the meantime, I am happy to pay for the development of the abovementioned engine that will enable MVP.
Objectives for the system:
1. It has to be completely offline (no online solutions will be accepted);
2. System has to be integrated with LLM (Llama or similar);
3. LLM must be capable on answering questions (via chatbot interface on GUI) on:
a. DP’s; and
b. Text it read while it was scanning the PDF and looking for DP’s;
c. Previously read documents related to the same client.
4. DP’s are to be saved to SQL table;
5. We need a dashboard, to train the engine on finding DP’s in the document.
Current problems:
A. Large part of documents are scanned, but regardless if the document is scanned or not, it’s all in digital form format. It just so happens that some forms were populated digitally, printed, and scanned;
B. As such, scanned documents are:
a. Tilted;
b. There is noise;
c. Zoom levels are different;
C. There are checkboxes;
D. DP’s can:
a. Change location within the document, but it will always have at least one anchor in fixed location relative to DP.
b. Few DP’s will have multiple entries that are true. In other words we will need unique ID per PDF per DP.
c. Some DP’s can be stretched over two pages.
THIS IS A LONG-TERM PROJECT. Project as defined above is to allow me to achieve MVP. Whoever creates me working engine, will become partner in business. In the meantime, I am happy to pay for the development of the abovementioned engine that will enable MVP.
Related categories:
Python
Software Architecture
Machine Learning (ML)
OCR
Large Language Models (LLMs)