Python Developer with Azure Ai services for building scalable Gen Ai Projects

Job ID: 39730771

Budget: ₹37,500 – ₹75,000 INR

Our organisation is ready to move the large-language-model work we have been prototyping into full production. I’d like an engineer who can take charge of everything from wiring Azure OpenAI into our on-prem / private-cloud stack to making sure the models run fast, safely, and cost-effectively once they are live.

What I need you to do
• Integrate the Azure-hosted GPT endpoints with several existing Python micro-services so our current applications can call them seamlessly.
• Tune throughput, latency, and token usage; set up the monitoring and alerting you’d expect in solid LLMOps (prompts, versioning, usage logs, rollback strategy, CI/CD pipelines).
• Extend the base models with custom features—prompt-engineering, embeddings, or fine-tuning—whenever a business unit has a new requirement.

I build mainly in Python, so your code, tests, and tooling should follow that ecosystem. Familiarity with FastAPI, Docker, and Kubernetes will help because they are already part of our pipeline. While Azure OpenAI is our primary platform, the ability to draw on other families such as native GPT-3/4, BERT, or T5 when the use-case demands would be a plus.

Deliverables (acceptance criteria)
– A production-ready LLM service running on our internal infrastructure, callable from our existing apps.
– RAG including NLP Queries with SQL for adding insights to data by asking questions
– Clean, well-commented Python code, unit tests, and step-by-step deployment documentation that let my team pick it up without hand-holding.

If this sounds like your wheelhouse, I’m looking forward to seeing how you’d approach the build-out and ongoing optimisation of our LLM stack.