AI/ML Engineer
Budget: ₹75,000 – ₹150,000 INR
We are building a next-generation AI-powered platform that combines legal & financial research, document drafting, and compliance automation.
As our AI/ML Engineer, you will design and implement retrieval-augmented generation (RAG) pipelines, optimize LLM performance, and build AI-powered drafting assistants. You will work closely with backend engineers, data engineers, and domain experts to ensure outputs are accurate, explainable, and citation-backed.
Key Responsibilities :
RAG Development :
- Implement semantic search over judgments, statutes, MCA/GST/RBI/SEBI data.
- Build chunking, embeddings, reranking, and citation-grounded retrieval pipelines.
LLM Orchestration:
- Integrate GPT-4/5, Claude, or self-hosted Llama/Mistral models.
- Apply guardrails to prevent hallucinations and enforce citation correctness.
Document Drafting AI:
- Develop AI-assisted template filling, clause recommender, and redlining systems.
- Support multi-lingual drafting (English + major Indian languages).
Data & Knowledge Graphs:
- Work with data engineers to preprocess legal/tax documents (OCR, NLP).
- Contribute to compliance knowledge graph (laws ↔ rules ↔ judgments).
Evaluation & Monitoring:
- Create golden datasets of legal/tax queries for model evaluation.
- Track metrics: Precision@1, citation accuracy, hallucination rate.
Collaboration:
- Partner with backend/frontend teams to expose AI APIs.
- Work with legal/CA domain experts to validate outputs.
Required Skills
- Strong programming in Python (NumPy, Pandas, PyTorch/TensorFlow).
- Experience with NLP/LLM frameworks: HuggingFace, LangChain, LlamaIndex.
- Knowledge of vector databases (Pinecone, Weaviate, Milvus, pgvector).
- Experience with RAG architectures (chunking, embeddings, reranking).
- Familiarity with information retrieval (BM25, hybrid search, cross-encoders).
- Deployment of LLMs (OpenAI/Anthropic APIs or open-source Llama/Mistral).
- Understanding of OCR/NLP preprocessing (Tesseract, SpaCy, FastText).-
- Good grasp of MLOps: Docker, API deployment, model serving (vLLM/TGI).
Preferred (Nice-to-Have)
- Prior work in legal tech, fintech, or compliance automation.
- Knowledge of Indian legal/tax datasets (Indian Kanoon, MCA21, CBDT/CBIC).
- Experience with multi-lingual NLP (Indic languages).
- Familiarity with knowledge graphs (Neo4j, RDF, SPARQL).
- Experience with evaluation frameworks (Ragas, DeepEval).
Education & Experience
- Bachelor’s or Master’s degree in Computer Science / Data Science / AI/ML (or equivalent experience).
- min. 5 years of experience in AI/ML/NLP projects (startup or research experience a plus).
As our AI/ML Engineer, you will design and implement retrieval-augmented generation (RAG) pipelines, optimize LLM performance, and build AI-powered drafting assistants. You will work closely with backend engineers, data engineers, and domain experts to ensure outputs are accurate, explainable, and citation-backed.
Key Responsibilities :
RAG Development :
- Implement semantic search over judgments, statutes, MCA/GST/RBI/SEBI data.
- Build chunking, embeddings, reranking, and citation-grounded retrieval pipelines.
LLM Orchestration:
- Integrate GPT-4/5, Claude, or self-hosted Llama/Mistral models.
- Apply guardrails to prevent hallucinations and enforce citation correctness.
Document Drafting AI:
- Develop AI-assisted template filling, clause recommender, and redlining systems.
- Support multi-lingual drafting (English + major Indian languages).
Data & Knowledge Graphs:
- Work with data engineers to preprocess legal/tax documents (OCR, NLP).
- Contribute to compliance knowledge graph (laws ↔ rules ↔ judgments).
Evaluation & Monitoring:
- Create golden datasets of legal/tax queries for model evaluation.
- Track metrics: Precision@1, citation accuracy, hallucination rate.
Collaboration:
- Partner with backend/frontend teams to expose AI APIs.
- Work with legal/CA domain experts to validate outputs.
Required Skills
- Strong programming in Python (NumPy, Pandas, PyTorch/TensorFlow).
- Experience with NLP/LLM frameworks: HuggingFace, LangChain, LlamaIndex.
- Knowledge of vector databases (Pinecone, Weaviate, Milvus, pgvector).
- Experience with RAG architectures (chunking, embeddings, reranking).
- Familiarity with information retrieval (BM25, hybrid search, cross-encoders).
- Deployment of LLMs (OpenAI/Anthropic APIs or open-source Llama/Mistral).
- Understanding of OCR/NLP preprocessing (Tesseract, SpaCy, FastText).-
- Good grasp of MLOps: Docker, API deployment, model serving (vLLM/TGI).
Preferred (Nice-to-Have)
- Prior work in legal tech, fintech, or compliance automation.
- Knowledge of Indian legal/tax datasets (Indian Kanoon, MCA21, CBDT/CBIC).
- Experience with multi-lingual NLP (Indic languages).
- Familiarity with knowledge graphs (Neo4j, RDF, SPARQL).
- Experience with evaluation frameworks (Ragas, DeepEval).
Education & Experience
- Bachelor’s or Master’s degree in Computer Science / Data Science / AI/ML (or equivalent experience).
- min. 5 years of experience in AI/ML/NLP projects (startup or research experience a plus).