Arabic NLP Annotation Platform

Job ID: 39749973

Budget: $8 – $15 USD

I’m building an end-to-end platform for collecting and labelling Arabic text—especially Saudi dialects—and I’d like a developer who can own the technical side. The very first milestone is getting a solid, Hugging Face–based pre-annotation pipeline in place; everything else (upload UI, review console, data export) will build on top of that core.

Here’s what I need implemented and wired together:

• A backend service (Python preferred) that calls the chosen Hugging Face model, receives raw sentences, and returns token-level or span-level labels ready for review.
• A simple web interface where users can drop in or bulk-upload text and immediately see the model’s suggestions, ready for correction in a human-in-the-loop workflow.
• Persistent storage of every annotation pass—raw text, model output, user-approved labels—in CSV, JSON, or a relational DB so the data is ready for downstream fine-tuning.
• Clean, documented code plus a read-me that lets my internal team spin the system up locally with Docker or a requirements file.

I’m comfortable with frameworks like FastAPI, Flask, or Streamlit for the UI layer, and PyTorch/Transformers under the hood. If you have a better stack in mind that still leverages Hugging Face models efficiently, let’s discuss—it just needs to handle Arabic encoding nuances and scale to thousands of samples without crawling.

Once this pre-annotation flow is solid we’ll iterate on advanced features—active learning loops, dialect detection, multi-label support—so I’m looking for someone who enjoys refining NLP products, not just delivering a quick proof of concept.

If you have hands-on experience shipping NLP tools for right-to-left languages, please share a brief note on similar projects and a quick outline of how you’d tackle the pre-annotation module. Looking forward to collaborating!