Build a Custom Arabic Legal SLM (Small Language Model) using HuggingFace + OCR + Entity Extraction

Job ID: 39513169

Budget: $250 – $750 USD

Project Title: AI Engineer Needed – Build Custom SLM with HuggingFace (OCR + NLP)

Description: We’re seeking an AI engineer experienced in building and fine-tuning a Small Language Model (SLM) using HuggingFace & PyTorch, tailored for Arabic legal workflows.

MVP Scope:

Load open-source SLM (e.g., Mistral, Phi-2)

Fine-tune on small legal dataset (commands, documents)

Implement OCR (e.g., Tesseract) to extract text from images (ID card, court notice)

Extract named entities (clients, dates, court types)

Generate sample legal petition in Word or text format

API or command-line interface acceptable

Delivery: within 10–14 days


Skills Required:

HuggingFace

PyTorch

OCR (Tesseract or similar)

NLP, entity extraction (Arabic legal context)

RPA or automation (Selenium/Puppeteer) – basic


Budget: $200–300
Timeline: 10–14 days
We will provide all training data in Excel and examples.

Please share:

Any GitHub or demo projects

Your experience with fine-tuning models or OCR in Arabic

Timeline estimate for this MVP