50-Page Viennese Street OCR
Budget: ₹12,500 – ₹37,500 INR
I have 50 high-resolution JPG/PNG pages that record Viennese street name changes—they map the historic to current ones. I need these pages converted into editable text with publication-grade accuracy. Because the files are already digital, no physical scanning is required; the job is all about precise OCR and meticulous proofreading.
You may use any robust engine you prefer—ABBYY FineReader, a finely-tuned Tesseract model, or similar—but the end result must be spotless German text, including every umlaut and ß exactly as printed. The layout is simple: two columns per entry, so please preserve that structure. Add any additional information and problems you encounter.
Deliverables
• A UTF-8 Excel or CSV file with columns: old_name | new_name | page_number
• A fully searchable PDF that mirrors the original pages for reference
Acceptance criteria
• ≥ 99 % character accuracy when compared to the source images
• All 55 pages accounted for, with no omissions or duplicates
• Correct diacritics and exact spellings for every street name
Once I verify a random sample against the originals and it meets these standards, the project is complete.
Its formatting rules are more complex than it seems. The only fixed thing is the 6-column structure: before after before after before after (except perhaps some shifts due to bad alignment of the scan)
Other than that a curly bracket shows which old names map to new names, the curly bracket can be on both sides. Basically the curly brackets determine how the page should be ocr-scanned deterministically. But note, that any approach to eager to rasterize row height seems to fall apart.
So output should be a mapping between old and new names and numbers, so multiple values to multiple values.
You will have to spot-check regularly to control the AI's output (I had the best results with Gemini, but it still was off occasionally, so I had to babysit it)
You may use any robust engine you prefer—ABBYY FineReader, a finely-tuned Tesseract model, or similar—but the end result must be spotless German text, including every umlaut and ß exactly as printed. The layout is simple: two columns per entry, so please preserve that structure. Add any additional information and problems you encounter.
Deliverables
• A UTF-8 Excel or CSV file with columns: old_name | new_name | page_number
• A fully searchable PDF that mirrors the original pages for reference
Acceptance criteria
• ≥ 99 % character accuracy when compared to the source images
• All 55 pages accounted for, with no omissions or duplicates
• Correct diacritics and exact spellings for every street name
Once I verify a random sample against the originals and it meets these standards, the project is complete.
Its formatting rules are more complex than it seems. The only fixed thing is the 6-column structure: before after before after before after (except perhaps some shifts due to bad alignment of the scan)
Other than that a curly bracket shows which old names map to new names, the curly bracket can be on both sides. Basically the curly brackets determine how the page should be ocr-scanned deterministically. But note, that any approach to eager to rasterize row height seems to fall apart.
So output should be a mapping between old and new names and numbers, so multiple values to multiple values.
You will have to spot-check regularly to control the AI's output (I had the best results with Gemini, but it still was off occasionally, so I had to babysit it)
Related categories:
Data Processing
Data Entry
Proofreading
Excel
PDF
OCR
German Translator
Text Recognition