OCR - Images and PDF Files
Budget: $30 – $250 USD
We need to build an OCR module that is capable of reading various image file extensions (eg jpg, png, gif, tiff, bmp), in addition to being able to read PDF files and identify images in their content.
1) The code must be able to read low resolution images, dark and images not centered on the page.
2) The code must be performative, as it will read thousands of files, totaling over 4 Terabytes
3) The solution must be able to be trained if necessary. Allowing fine-tuning if necessary. Initially we will work with Portuguese/Brazil (PT-BR)
4) The code must be made available in an open C# solution
5) The solution should export the results to PLAIN TEXT.
6) Delivery of the solution: We expect a project in Visual Studio developed in C#, where we can test all the features above. A simple web screen or forms will be ok. Also an explanation of how we can train the system.
We will provide some files for tests and examples.
1) The code must be able to read low resolution images, dark and images not centered on the page.
2) The code must be performative, as it will read thousands of files, totaling over 4 Terabytes
3) The solution must be able to be trained if necessary. Allowing fine-tuning if necessary. Initially we will work with Portuguese/Brazil (PT-BR)
4) The code must be made available in an open C# solution
5) The solution should export the results to PLAIN TEXT.
6) Delivery of the solution: We expect a project in Visual Studio developed in C#, where we can test all the features above. A simple web screen or forms will be ok. Also an explanation of how we can train the system.
We will provide some files for tests and examples.