Build a RAG-Based AI System for a Large Arabic Text Corpus

Job ID: 40410216

Budget: $750 – $1,500 USD

We are looking for a highly experienced AI/LLM engineer or technical team to build a proof of concept for a private Arabic content platform with a very large structured Arabic text corpus.

This is NOT a basic chatbot project.

We need a technically strong RAG-based system that can search, retrieve, reason over, and generate source-grounded answers from our own Arabic database.

The system must not rely on the LLM’s general memory. It must retrieve relevant records from our own corpus and generate answers strictly based on retrieved context, with clear source references.

Required technical scope:

- RAG architecture design
- Arabic text preprocessing and normalization
- Embeddings strategy for Arabic text
- Vector database setup
- Hybrid search: semantic search + keyword/BM25 search
- Metadata filtering
- Source-grounded answer generation
- Hallucination reduction and answer validation
- LLM API integration: OpenAI, Claude, Gemini, or similar
- Backend API for integration with an existing web platform
- Evaluation framework to measure retrieval quality and answer accuracy

Preferred tools and technologies:

- Python
- FastAPI or similar backend framework
- LangChain or LlamaIndex
- Pinecone, Qdrant, Weaviate, pgvector, or similar
- Elasticsearch / OpenSearch / BM25
- OpenAI API / Anthropic Claude API / Gemini API
- Arabic NLP experience
- Experience with large text datasets

Initial deliverable:

We want to start with a limited Proof of Concept using a sample dataset, not the full database.

The PoC should demonstrate:

1. Data ingestion and indexing
2. Arabic text normalization
3. Embeddings and vector search
4. Hybrid retrieval
5. Source-linked answers
6. A simple API or demo interface
7. Evaluation results for retrieval and answer quality

Important:

Please do not apply if your experience is limited to building simple ChatGPT wrappers or generic website chatbots.

In your proposal, you must answer the following:

1. Describe your proposed RAG architecture.
2. Which vector database would you choose and why?
3. How would you handle a large Arabic text corpus?
4. How would you prevent hallucinated answers?
5. How would you combine semantic search with keyword search?
6. How would you evaluate answer accuracy?
7. Share one real RAG / semantic search / LLM project you have built.

Project details, data samples, and platform identity will be shared only with shortlisted candidates.