YOLO Optimisation for Scalable Object Detection
Budget: ₹1,500 – ₹12,500 INR
Description:
We are running a YOLO-based object detection system deployed on a Google Cloud Platform (GCP) Kubernetes setup with a worker-handler architecture. Each node is currently assigned two CPU cores.
During normal operations, when 50 users are active, CPU usage peaks around 1200 millicores (mc), raising concerns about scalability. We aim to support up to 15,000 concurrent users, with the goal of keeping CPU usage under 500mc per instance, without compromising the accuracy and reliability of detections.
Project Goals:
We are seeking an experienced AI/Machine Learning/DevOps engineer to analyze our current system and implement optimizations. The goal is to reduce CPU load while maintaining accurate object detection and classifications.
Key Objectives:
Analyze Current Architecture:
Review YOLO model version (v3/v4/v5/v8 or custom).
Understand how it's integrated within the worker-handler setup.
Evaluate CPU usage patterns and identify performance bottlenecks.
Propose and Implement Optimizations:
Suggest alternatives or optimizations (e.g., quantization, pruning, TensorRT conversion, ONNX usage, lighter model variants).
Propose architectural improvements (batching requests, async processing, load balancing).
Implement changes that reduce CPU usage and improve performance at scale.
Scalability Validation:
Simulate or benchmark performance under higher loads (up to 15,000 users).
Ensure consistent detection performance and acceptable latency.
Deployment and Testing:
Help with optimized deployment in GCP Kubernetes.
Monitor improvements using metrics (CPU, latency, detection accuracy).
We are running a YOLO-based object detection system deployed on a Google Cloud Platform (GCP) Kubernetes setup with a worker-handler architecture. Each node is currently assigned two CPU cores.
During normal operations, when 50 users are active, CPU usage peaks around 1200 millicores (mc), raising concerns about scalability. We aim to support up to 15,000 concurrent users, with the goal of keeping CPU usage under 500mc per instance, without compromising the accuracy and reliability of detections.
Project Goals:
We are seeking an experienced AI/Machine Learning/DevOps engineer to analyze our current system and implement optimizations. The goal is to reduce CPU load while maintaining accurate object detection and classifications.
Key Objectives:
Analyze Current Architecture:
Review YOLO model version (v3/v4/v5/v8 or custom).
Understand how it's integrated within the worker-handler setup.
Evaluate CPU usage patterns and identify performance bottlenecks.
Propose and Implement Optimizations:
Suggest alternatives or optimizations (e.g., quantization, pruning, TensorRT conversion, ONNX usage, lighter model variants).
Propose architectural improvements (batching requests, async processing, load balancing).
Implement changes that reduce CPU usage and improve performance at scale.
Scalability Validation:
Simulate or benchmark performance under higher loads (up to 15,000 users).
Ensure consistent detection performance and acceptable latency.
Deployment and Testing:
Help with optimized deployment in GCP Kubernetes.
Monitor improvements using metrics (CPU, latency, detection accuracy).
Related categories:
Machine Learning (ML)
Infrastructure Architecture
AI (Artificial Intelligence) HW/SW