Hindi POS Tagging for PDFs
Budget: ₹100 – ₹400 INR
I have a collection of Hindi-language documents supplied as PDFs that need to be annotated for Part-of-Speech. Your task is to extract the text from each PDF and tag every token with the correct POS label, following the standard Hindi tag set (for example, NN, VB, JJ, PRP, etc.).
Accuracy is essential because the data will feed a downstream NLP model. If you already use tools such as spaCy, Stanza, or custom tagging interfaces, feel free to integrate them, but the final output must be a clean, human-checked file—not just an auto-tagged draft.
Deliverables
• For each input PDF: a UTF-8 text file containing the original sentence order alongside its POS tags (tab-separated or CoNLL-style).
• A brief log noting any unreadable sections or encoding issues you encounter.
I will share the PDFs and a short style guide once we start. Let me know your estimated turnaround time per 1,000 words and any previous work you’ve done with Hindi linguistic annotation.
Accuracy is essential because the data will feed a downstream NLP model. If you already use tools such as spaCy, Stanza, or custom tagging interfaces, feel free to integrate them, but the final output must be a clean, human-checked file—not just an auto-tagged draft.
Deliverables
• For each input PDF: a UTF-8 text file containing the original sentence order alongside its POS tags (tab-separated or CoNLL-style).
• A brief log noting any unreadable sections or encoding issues you encounter.
I will share the PDFs and a short style guide once we start. Let me know your estimated turnaround time per 1,000 words and any previous work you’ve done with Hindi linguistic annotation.