Multilingual AI on Proprietary Data
Budget: ₹1,500 – ₹12,500 INR
I need a ChatGPT-style assistant that can read every file my organisation generates—PDFs, Word docs, Excel sheets, HTML pages and anything else we store—then answer questions using that private knowledge base. The assistant must recognise the language of the query and reply in the same tongue; Hindi, Marathi, English and Telugu are the priority.
The job covers the full pipeline: ingesting and parsing documents, building a semantic index or vector store, connecting it to a generative model and exposing everything through a simple chat interface. Whichever mix of OpenAI, open-source LLMs, embeddings or retrieval-augmented generation you choose is up to you, as long as the final experience feels as fluid as ChatGPT while keeping my data secure on our own servers.
Please send a detailed project proposal describing
• how you will extract and normalise content from the various formats
• your multilingual strategy (tokenisers, embeddings, fine-tuning, or any other approach)
• the tech stack for search, orchestration and the front-end chat window
• milestones, delivery timeline and post-delivery support
Deliverables
1. Automated pipeline that continuously ingests new documents and updates the index
2. Multilingual Q&A model wired to that index, with language auto-detection and same-language response
3. Web-based chat UI (desktop and mobile friendly)
4. Deployment guide and source code
Acceptance will be based on accurate answers drawn from my files and correct language handling across all four languages.
The job covers the full pipeline: ingesting and parsing documents, building a semantic index or vector store, connecting it to a generative model and exposing everything through a simple chat interface. Whichever mix of OpenAI, open-source LLMs, embeddings or retrieval-augmented generation you choose is up to you, as long as the final experience feels as fluid as ChatGPT while keeping my data secure on our own servers.
Please send a detailed project proposal describing
• how you will extract and normalise content from the various formats
• your multilingual strategy (tokenisers, embeddings, fine-tuning, or any other approach)
• the tech stack for search, orchestration and the front-end chat window
• milestones, delivery timeline and post-delivery support
Deliverables
1. Automated pipeline that continuously ingests new documents and updates the index
2. Multilingual Q&A model wired to that index, with language auto-detection and same-language response
3. Web-based chat UI (desktop and mobile friendly)
4. Deployment guide and source code
Acceptance will be based on accurate answers drawn from my files and correct language handling across all four languages.