Python ML Engineer for PDF Data Extraction
Budget: €250 – €750 EUR
Dear developers,
I am looking for a skilled python engineer that is able to create a ML engine which is capable to extract specific datapoints from the pdf files and save them as per set of rules to sql table.
system has to run offline, so no online solutions will be accepted as they cannot be used for this project.
current problems:
1. large part of documents are scanned, but data is digital. essentially form was printed and scanned upon completion.
2. scanned files are tilted.
3. scanned files have noise which require removal of it.
4. few scanned documents require quality enhancement.
5. some scanned documents have different zoom.
6. all files have checkboxes.
7. specific datapoints can change location within the file, so engine must search for it and find it. there is master list at front page that indicates coa checkboxes as to what data points will be in the file.
8. few datapoints will have multiple entries, that can be true and all instances have to be entered into SQL segregating them by unique identifiers.
we also need dashboard, that allows training of OCR engine using ML such as llama.
to repeat, system has to run completely offline.
bonus task: integrated ML that is able to scan the documents per group, and combined with extracted data is able to answer questions about the text within such documents, per group, bot mixing up information with other groups.
I am looking for a skilled python engineer that is able to create a ML engine which is capable to extract specific datapoints from the pdf files and save them as per set of rules to sql table.
system has to run offline, so no online solutions will be accepted as they cannot be used for this project.
current problems:
1. large part of documents are scanned, but data is digital. essentially form was printed and scanned upon completion.
2. scanned files are tilted.
3. scanned files have noise which require removal of it.
4. few scanned documents require quality enhancement.
5. some scanned documents have different zoom.
6. all files have checkboxes.
7. specific datapoints can change location within the file, so engine must search for it and find it. there is master list at front page that indicates coa checkboxes as to what data points will be in the file.
8. few datapoints will have multiple entries, that can be true and all instances have to be entered into SQL segregating them by unique identifiers.
we also need dashboard, that allows training of OCR engine using ML such as llama.
to repeat, system has to run completely offline.
bonus task: integrated ML that is able to scan the documents per group, and combined with extracted data is able to answer questions about the text within such documents, per group, bot mixing up information with other groups.