Senior AI Engineer for Voice-Ordering System

Job ID: 40248359

Budget: $250 – $750 USD

I am looking for a senior AI/Backend Engineer to build a high-performance voice-ordering system for restaurants. The system must handle "Code-switching" (mixed Arabic and Hebrew dialects) with ultra-low latency.

Key Requirements:

STT Engine: Implement Faster-Whisper (Large-v3-Turbo) on a GPU-accelerated environment (RunPod/AWS/Lambda Labs).

Contextual Accuracy: Use Initial Prompting techniques to ensure high accuracy for specific restaurant menus (dishes, modifiers, doneness levels).

Real-time Processing: Experience with WebSockets for streaming audio and achieving <1s end-to-end latency.

Parsing: Integrate an LLM (GPT-4o-mini or local Llama-3) to convert raw transcribed text into structured JSON.

Infrastructure: Knowledge of Serverless GPUs or Dockerized GPU deployments to ensure scalability for 100+ concurrent restaurants.

Specific Challenges:
The system must be optimized for "Arabic-Hebrew" hybrid speech (Local dialect). Experience with VAD (Voice Activity Detection) like Silero is a plus to handle background noise in kitchens.

Deliverables:

A working backend API that accepts audio and returns structured JSON.

Latency optimization report.

Scalability plan for multi-tenant architecture.