Senior LLM Infrastructure and Backend DevOps Engineer
Budget: $50 – $0 USD
Summary
We are seeking a Senior LLM Infrastructure and DevOps Engineer to architect, deploy, and maintain large scale generative AI systems across AWS and GCP. This role focuses on LLM inference pipelines, multi agent orchestration, infra as code, and cloud native automation. You will own the reliability, security, and scalability of all infrastructure supporting N1’s AI driven report generation platform.
This is an advanced role for an engineer deeply familiar with deploying production LLM workflows, optimizing cloud spend for generative AI, and implementing secure, highly available multi cloud systems.
Core Requirements
• Infrastructure diagramming and architecture ownership
• 8 plus years in DevOps or CloudOps, including containerized and distributed deployments
• Expert in Terraform, GitHub Actions, AWS CDK, and GCP deployment workflows
• Strong experience with AWS CodePipelines
• Full management of AWS services (IAM, ECS or Fargate, S3, Bedrock)
• Strong experience with GCP services (Cloud Run, Batch with batch job management, Firestore)
• Advanced Docker, CI or CD, and automated build and deploy pipelines
• Proven experience deploying LLM systems at scale, including Bedrock, Claude, Gemini, Hugging Face, and others via LiteLLM
• Security expertise: IAM policies, secrets management, audit logging, intrusion detection
• Deep experience with cost control for LLM inference workloads
• Experience with AWS CloudWatch, Cost Explorer, and model usage tracking
• Ability to implement model spend monitoring (LiteLLM or equivalent)
• Familiarity with prompt management systems like Bedrock Prompt Management
• Ability to enforce provider specific quotas (Gemini token limits, request quotas, etc.)
• Ability to integrate new model endpoints including Imagen 3, Google Multimodal, grok 3 beta
• Strong debugging skills for LLM inference failures, image generation issues, and vector fallback logic (Weaviate, ChromaDB)
Additional Required Experience
• Fluent in switching and deploying across LLM providers (OpenAI, Google, Anthropic, DeepSeek)
• Experience deploying secure HTTP or HTTPS transports and server sent events
• Ability to diagnose failures across distributed file processing and LLM extraction pipelines
• Proficiency with multimodal API constraints and data governance requirements
• Experience handling rapid model deprecations or replacements
• Strong cross functional collaboration with backend and data engineering teams
Deliverables
Complete infrastructure diagrams for AWS and GCP environments
Production grade LLM inference pipelines with automated CI or CD
Fully coded Terraform or CDK infrastructure modules
Stable and optimized multi agent LLM orchestration deployed across providers
Model cost monitoring dashboards and LLM usage reporting
Logging and monitoring improvements for distributed extraction pipelines
Robust fallback logic for vector database layers (Weaviate, ChromaDB)
Secure IAM configurations and documented secret management workflows
Automated model quota enforcement and routing logic
Documentation for deployment processes, rollback plans, and disaster recovery
Cost reduction strategies validated in production
Stable cross cloud collaboration pipelines with backend or data teams
We are seeking a Senior LLM Infrastructure and DevOps Engineer to architect, deploy, and maintain large scale generative AI systems across AWS and GCP. This role focuses on LLM inference pipelines, multi agent orchestration, infra as code, and cloud native automation. You will own the reliability, security, and scalability of all infrastructure supporting N1’s AI driven report generation platform.
This is an advanced role for an engineer deeply familiar with deploying production LLM workflows, optimizing cloud spend for generative AI, and implementing secure, highly available multi cloud systems.
Core Requirements
• Infrastructure diagramming and architecture ownership
• 8 plus years in DevOps or CloudOps, including containerized and distributed deployments
• Expert in Terraform, GitHub Actions, AWS CDK, and GCP deployment workflows
• Strong experience with AWS CodePipelines
• Full management of AWS services (IAM, ECS or Fargate, S3, Bedrock)
• Strong experience with GCP services (Cloud Run, Batch with batch job management, Firestore)
• Advanced Docker, CI or CD, and automated build and deploy pipelines
• Proven experience deploying LLM systems at scale, including Bedrock, Claude, Gemini, Hugging Face, and others via LiteLLM
• Security expertise: IAM policies, secrets management, audit logging, intrusion detection
• Deep experience with cost control for LLM inference workloads
• Experience with AWS CloudWatch, Cost Explorer, and model usage tracking
• Ability to implement model spend monitoring (LiteLLM or equivalent)
• Familiarity with prompt management systems like Bedrock Prompt Management
• Ability to enforce provider specific quotas (Gemini token limits, request quotas, etc.)
• Ability to integrate new model endpoints including Imagen 3, Google Multimodal, grok 3 beta
• Strong debugging skills for LLM inference failures, image generation issues, and vector fallback logic (Weaviate, ChromaDB)
Additional Required Experience
• Fluent in switching and deploying across LLM providers (OpenAI, Google, Anthropic, DeepSeek)
• Experience deploying secure HTTP or HTTPS transports and server sent events
• Ability to diagnose failures across distributed file processing and LLM extraction pipelines
• Proficiency with multimodal API constraints and data governance requirements
• Experience handling rapid model deprecations or replacements
• Strong cross functional collaboration with backend and data engineering teams
Deliverables
Complete infrastructure diagrams for AWS and GCP environments
Production grade LLM inference pipelines with automated CI or CD
Fully coded Terraform or CDK infrastructure modules
Stable and optimized multi agent LLM orchestration deployed across providers
Model cost monitoring dashboards and LLM usage reporting
Logging and monitoring improvements for distributed extraction pipelines
Robust fallback logic for vector database layers (Weaviate, ChromaDB)
Secure IAM configurations and documented secret management workflows
Automated model quota enforcement and routing logic
Documentation for deployment processes, rollback plans, and disaster recovery
Cost reduction strategies validated in production
Stable cross cloud collaboration pipelines with backend or data teams