Multi-Camera Kitchen AI System Development

Job ID: 40540089

Budget: $250 – $750 USD

## Job Title
Computer Vision Engineer: Custom Multi-Camera Kitchen Automation System (YOLO + Cloud Sync)
## Job Description## Project Overview
We operate a commercial kitchen with 8 cooking stations and are looking to build a custom, hands-free quality control system. The goal is to automatically capture a 3-second video clip every time a cook adds a new ingredient into a cooking pot or pan.
To make this highly accurate and lightweight, we are standardizing our prep containers. Cooks will transfer ingredients into uniform, highly visible, color-coded prep bowls. The AI needs to track these specific containers, detect when they hover and tilt over a cooking zone, crop a 3-second video clip (1 second before the tilt, 2 seconds after), and upload it to a cloud dashboard.
## System Architecture & Hardware

* Inputs: 8 Overhead PoE (Power over Ethernet) Dome Cameras mounted directly above 8 separate cooking stations.
* Local Compute: 1 Central Edge AI Server/PC (equipped with a high-end NVIDIA GPU, e.g., RTX 4000 series) processing the 8 live streams simultaneously.
* Output: Short 3-second compressed video clips sent directly to a cloud storage bucket.
* Frontend: A simple, mobile-friendly web dashboard where management can view clips sorted by Station # and Time/Date.

## Key Responsibilities

* Set up the multi-camera RTSP streaming pipeline to the central local processing machine.
* Develop and train an object detection/action segmentation model (preferably using YOLOv8/YOLOv10 or similar light models) to identify our specific prep bowls and detect the "tilting/pouring" action.
* Implement a rolling video buffer logic that extracts exactly 3 seconds of video surrounding the detected action trigger.
* Build the Edge-to-Cloud pipeline to upload compressed video fragments efficiently to a secure cloud platform (AWS S3, Google Cloud, or Azure).
* Build a clean, lightweight frontend interface (using tools like Retool, Streamlit, or a basic React app) for cloud video playback.

## Required Skills

* Deep expertise in Computer Vision and Deep Learning (Python, OpenCV, PyTorch/TensorFlow).
* Extensive experience with real-time object detection models (YOLO workflow is highly preferred).
* Proven track record building multi-camera RTSP video pipelines and handling edge processing hardware (NVIDIA Jetson or GPU workstations).
* Cloud architecture experience (AWS S3 / Lambda or Google Cloud equivalents).
* Experience with video encoding and compression (FFmpeg) to minimize cloud storage costs.

## Project Type & Budget

* Project Type: One-time project with potential for ongoing retainer/maintenance contracts.
* Budget: Flexible / Competitive based on experience and proposed timeline.
* Please include links to any previous computer vision or video-processing projects you have built in your proposal.