Expert for Arabic PDF Conversion & AI_Search
Budget: $30 – $250 USD
Arabic PDF Data Structuring & AI Search Specialist
We are looking for an experienced freelancer or full-time specialist to convert one chapter from an Arabic PDF book into structured, searchable data.
This is a Proof of Concept on one chapter only, not a full-book project at this stage.
The task includes:
Arabic text extraction.
Arabic OCR cleanup.
Mixed Arabic/English text handling.
PDF layout analysis.
Image extraction.
Table extraction.
Content chunking.
JSON schema creation.
Concept extraction.
Question/exercise extraction, if available.
Page-level source referencing.
Preparing the data for semantic search, vector search, and RAG systems.
Providing documentation and a quality report.
Required experience:
Previous work with Arabic PDF content.
Arabic OCR.
Python.
PDF processing.
JSON data modeling.
Search-ready data preparation.
Embeddings, semantic search, or RAG experience preferred.
Deliverables:
Structured JSON files.
Extracted images and tables.
Search-ready chunks.
Sample queries or a simple demo.
Methodology documentation.
Quality report.
Please apply with:
Previous Arabic PDF/OCR examples.
Tools you will use.
Timeline.
Cost.
Sample JSON schema.
Explanation of your approach.
Important: This is only a test project for one chapter from one Arabic book. A larger project may be discussed later depending on the quality of the output.
We are looking for an experienced freelancer or full-time specialist to convert one chapter from an Arabic PDF book into structured, searchable data.
This is a Proof of Concept on one chapter only, not a full-book project at this stage.
The task includes:
Arabic text extraction.
Arabic OCR cleanup.
Mixed Arabic/English text handling.
PDF layout analysis.
Image extraction.
Table extraction.
Content chunking.
JSON schema creation.
Concept extraction.
Question/exercise extraction, if available.
Page-level source referencing.
Preparing the data for semantic search, vector search, and RAG systems.
Providing documentation and a quality report.
Required experience:
Previous work with Arabic PDF content.
Arabic OCR.
Python.
PDF processing.
JSON data modeling.
Search-ready data preparation.
Embeddings, semantic search, or RAG experience preferred.
Deliverables:
Structured JSON files.
Extracted images and tables.
Search-ready chunks.
Sample queries or a simple demo.
Methodology documentation.
Quality report.
Please apply with:
Previous Arabic PDF/OCR examples.
Tools you will use.
Timeline.
Cost.
Sample JSON schema.
Explanation of your approach.
Important: This is only a test project for one chapter from one Arabic book. A larger project may be discussed later depending on the quality of the output.