Local AI LLM + RAG System Development

Job ID: 39364060

Budget: $3,000 – $5,000 USD

Key Requirements
• Ingest, clean, and index our text, PDF, and PowerPoint files, plus proprietary chart JSON data (from a local API)
• Fine-tune or supervise-train an open-source LLM (Llama 3, Mistral, or similar) on this data
• Implement a retrieval-augmented generation (RAG) pipeline with a local vector database (FAISS/ChromaDB)
• All data, models, and inference must remain strictly local—no cloud, no external API calls
• Expose a robust and secure REST API for querying the system, designed for easy integration with future apps and reporting tools (e.g., JasperReports)
• Provide user-friendly tools (CLI or basic web interface) for non-developer ingestion of new data, with full documentation
• Deploy, test, and document everything on our local GPU server, and support deployment to additional production servers
• Provide clear, detailed documentation for setup, usage, retraining, and future integration

Deliverables
• Fully functional local LLM+RAG system
• REST API with documentation
• Data ingestion tools and scripts
• All source code and configuration
• Deployment and handover support
• Documentation for IT and user teams

Timeline
We expect the project to be completed in 6–8 weeks, with multiple tasks running in parallel. See our attached requirements and project timeline table for details.

Skills & Experience
• LLM/RAG implementation (HuggingFace, PyTorch/TensorFlow)
• Data extraction from unstructured and JSON sources
• API development (FastAPI/Flask)
• Local deployment on Windows and/or Linux GPU servers
• Strong documentation and communication skills