Custom AI Model Build & Deploy

Job ID: 39809415

Budget: $8 – $15 USD

The Applied ML Engineer (Gen AI) plays a foundational role in establishing and scaling the organization’s AI infrastructure and knowledge systems. This position focuses on building robust LLM deployment pipelines, retrieval-augmented generation (RAG) systems, embedding/tokenization workflows, and integrating AI agents into internal tools. The engineer will collaborate across teams to develop production-ready Gen AI applications and enable internal innovation at scale. The ideal candidate combines technical depth in machine learning operations with creativity in applying AI to real-world business contexts.

How You’ll Make an Impact

LLM Infrastructure & Evaluation
• Deploy, manage, and monitor Large Language Models (LLMs) such as LLaMA, GPT, Mistral, and Claude.
• Build internal evaluation pipelines to benchmark model performance across diverse business use cases.

Knowledge Systems Development
• Design and implement advanced RAG pipelines using vector databases like Pinecone or Weaviate.
• Develop graph-powered knowledge systems using Neo4j, ArangoDB, or TigerGraph.
• Create scalable chunking and tokenization strategies for semantic search and document ingestion.

Embedding, Tokenization & Retrieval Optimization
• Select and apply embedding models (OpenAI, Cohere, HuggingFace) for downstream tasks.
• Optimize chunking and retrieval pipelines for low-latency, high-recall performance.
• Engineer dynamic context windows for multi-turn conversations and long-form reasoning.

Agentic Workflows & Tooling Integration
• Build autonomous workflows using LangChain, Haystack, AutoGen, or CrewAI.
• Integrate LLM tools into internal platforms (e.g., task bots, report generators, coding copilots).
• Automate tool use via Function Calling APIs and connect agents to knowledge graphs.

MLOps Core & Deployment
• Build CI/CD pipelines for data workflows, LLM retraining, and production rollout.
• Manage containerized environments using Docker, Kubernetes, or Kubeflow.
• Implement monitoring, alerting, and logging for Gen AI systems in production.

Experimentation & Internal Enablement
• Lead experiments and PoCs to evaluate new LLM use cases (chatbots, assistants, research agents).
• Collaborate with backend, data, and QA teams to scale AI solutions company-wide.
• Create reusable AI modules to support operational teams (e.g. Compliance, Finance, Payouts).

What You Bring
• Bachelor’s degree in Computer Science, AI, Data Engineering, or related field.
• 2+ years of experience in Machine Learning, MLOps, or AI infrastructure.
• Proficiency in deploying and optimizing LLMs and vector databases.
• Hands-on experience with LangChain, LlamaIndex, or similar frameworks.
• Strong understanding of tokenization, embeddings, and prompt engineering.
• Experience with graph databases (e.g., Neo4j, ArangoDB) and query languages (Cypher, GraphQL).
• Familiarity with fine-tuning methods (LoRA, QLoRA) and feedback loops.
• Containerization and orchestration experience with Docker/Kubernetes.
• Team mindset, adaptability, and eagerness to solve complex problems in novel ways.

Your X-Factor
• Designs AI infrastructure that’s production-ready, not just experimental.
• Automates LLM workflows using scripting and orchestration frameworks.
• Balances innovation with maintainability in a fast-scaling environment.
• Identifies opportunities to embed AI into core business operations.
• Bridges technical depth with clear communication and business alignment.