Data Entry for Machine Learning
Budget: €8 – €30 EUR
I’m preparing a new machine-learning model and need the raw text fed in quickly and accurately. All of the material lives in scanned PDFs—no other formats—and each page must be turned into clean, plain UTF-8 text.
Here’s what the work looks like: you’ll open every PDF, extract or carefully re-type the content, and proof the result so line breaks, punctuation, and any special characters match the source exactly. The final text should be free of headers, footers, or artefacts that could confuse downstream preprocessing.
Deliverables
• One .txt file per original PDF, named identically
• A simple spreadsheet listing: file name, page count, word count, completion timestamp
Acceptance criteria
• 100 % coverage of the visible text in every PDF supplied
• ≤ 0.5 % character-level error rate when spot-checked against the scans
• Consistent encoding (UTF-8) and no hidden formatting codes
If you already work with tools such as Adobe Acrobat OCR, Tesseract, or Abbyy FineReader that’s great—just be ready to proofread and correct the inevitable OCR slips. Speed helps, but accuracy is king; let me know your approximate throughput in pages per hour and how soon you can start.
Here’s what the work looks like: you’ll open every PDF, extract or carefully re-type the content, and proof the result so line breaks, punctuation, and any special characters match the source exactly. The final text should be free of headers, footers, or artefacts that could confuse downstream preprocessing.
Deliverables
• One .txt file per original PDF, named identically
• A simple spreadsheet listing: file name, page count, word count, completion timestamp
Acceptance criteria
• 100 % coverage of the visible text in every PDF supplied
• ≤ 0.5 % character-level error rate when spot-checked against the scans
• Consistent encoding (UTF-8) and no hidden formatting codes
If you already work with tools such as Adobe Acrobat OCR, Tesseract, or Abbyy FineReader that’s great—just be ready to proofread and correct the inevitable OCR slips. Speed helps, but accuracy is king; let me know your approximate throughput in pages per hour and how soon you can start.
Related categories:
Data Processing
Data Entry
Proofreading
PDF
OCR
Data Cleansing
Data Extraction
Data Analysis
Adobe Acrobat
Data Management