Tutor for tesseract font training [OCR]
Budget: ₹2,500 – ₹3,500 INR
I have a concrete training case in tesseract, I'm stuck and a little support would be helpful.
Specifically, I i need OCR for PDFs scanned from ASCII text files using the font "Courier new" (So monospace, serif)
With the existing traindat eng or deu I get rates of around 99%, but I would like to increase the rate further, because
* it is always the same font
* we have a reduced charset
* usually the same letters are recognized incorrectly or are ignored.
I already have a setup ~20 one-page PDFs with the "ground truth", and generated .tif and .box files.
Already tried to use tesseract console tools, jTextBoxEditor, TessTrainGUI 6.4;
i tried to make new font traing and based on tessdata_best/deu.traineddata.
No success.
YOUR Job:
Just analyze the problem (you get sample files from me) and give me some hints, what steps to do to get improved bets possible traineddata set.
Specifically, I i need OCR for PDFs scanned from ASCII text files using the font "Courier new" (So monospace, serif)
With the existing traindat eng or deu I get rates of around 99%, but I would like to increase the rate further, because
* it is always the same font
* we have a reduced charset
* usually the same letters are recognized incorrectly or are ignored.
I already have a setup ~20 one-page PDFs with the "ground truth", and generated .tif and .box files.
Already tried to use tesseract console tools, jTextBoxEditor, TessTrainGUI 6.4;
i tried to make new font traing and based on tessdata_best/deu.traineddata.
No success.
YOUR Job:
Just analyze the problem (you get sample files from me) and give me some hints, what steps to do to get improved bets possible traineddata set.