Senior AI Backend Engineer Needed
Budget: $3,000 – $5,000 USD
Our product relies on Large Language Models, and I need a seasoned backend engineer to turn those models into a rock-solid, production-ready service. Your main responsibility is API development: designing, building, and optimising Python-based endpoints that expose LLM features to our web and mobile clients.
You should feel at home writing clean, test-covered code in FastAPI (or a comparable framework) and managing everything that makes an API reliable—auth layers, rate-limiting, logging, CI/CD, and containerisation. Although the core models are LLMs, familiarity with common AI stacks such as TensorFlow, PyTorch or Scikit-Learn will help when we experiment with alternative architectures.
Key deliverables
• A version-controlled codebase in Python that wraps our existing LLM checkpoints behind REST (or gRPC) endpoints
• Dockerfile and deployment scripts for staging and production
• Unit and integration tests with ≥90 % coverage, plus concise Swagger/OpenAPI docs
• Monitoring hooks (Prometheus/Grafana or similar) so we can track uptime and latency
Once the first milestone is live we will iterate together on performance tuning, caching, and scaling strategies, so a proactive approach to profiling and optimisation is essential.
If building high-performance AI APIs excites you, let’s talk timelines and dive straight into the repo.
You should feel at home writing clean, test-covered code in FastAPI (or a comparable framework) and managing everything that makes an API reliable—auth layers, rate-limiting, logging, CI/CD, and containerisation. Although the core models are LLMs, familiarity with common AI stacks such as TensorFlow, PyTorch or Scikit-Learn will help when we experiment with alternative architectures.
Key deliverables
• A version-controlled codebase in Python that wraps our existing LLM checkpoints behind REST (or gRPC) endpoints
• Dockerfile and deployment scripts for staging and production
• Unit and integration tests with ≥90 % coverage, plus concise Swagger/OpenAPI docs
• Monitoring hooks (Prometheus/Grafana or similar) so we can track uptime and latency
Once the first milestone is live we will iterate together on performance tuning, caching, and scaling strategies, so a proactive approach to profiling and optimisation is essential.
If building high-performance AI APIs excites you, let’s talk timelines and dive straight into the repo.