AI Voice System Configuration & Fine-Tuning

Job ID: 40206286

Budget: ₹12,500 – ₹37,500 INR

We are building a production-grade AI Calling Agent using a 2-VM architecture:

VM-1 → Telephony, SIP trunk, media handling, STT & TTS (THIS TASK)

VM-2 → AI Brain (LLM + Vector DB) – handled separately

This project is strictly limited to VM-1.
The freelancer will set up telephony, real-time audio processing, STT/TTS, API connectivity, and perform full fine-tuning so that live calls sound natural, stable, and low-latency.

Scope of Work (VM-1 Only)
VM & Network Setup

OS: Ubuntu (preferred) on Proxmox

Configure:

Public interface → SIP + RTP

Private interface → AI Brain API

Optimize OS for:

Low-latency audio

High concurrent calls

Firewall configuration:

SIP (TLS)

RTP port range

Internal API access only

SIP Trunk & Telephony Configuration

Configure SIP Trunk (provider details will be shared)

SIP over TLS preferred

Handle:

Incoming & outgoing calls

NAT traversal

DTMF

Call start / end events

Codec setup:

Opus (primary)

G.711 (fallback)

Telephony stack:

FreeSWITCH / Kamailio / OpenSIPS (freelancer must justify choice)

Media & Call Flow Handling

Handle RTP audio streams in real time

Call flow:

Receive caller audio

Stream audio to STT

Receive AI response

Convert to TTS

Play back audio to caller

Support:

Voice Activity Detection (VAD)

Silence detection

Caller barge-in (interrupt AI speech)

STT (Speech-to-Text) Integration

Integrate real-time STT

Audio streaming in chunks

Handle:

Partial transcripts

Final transcripts

Tune for:

Indian accents

Background noise

Fast & slow speakers

Ensure no loss of first/last words

TTS (Text-to-Speech) Integration

Integrate low-latency TTS

Telephony-optimized voices

Tune:

Speed

Pitch

Volume normalization

Playback must:

Start quickly

Sound natural

Stop immediately on barge-in

API Layer for AI Brain Connectivity

Expose secure APIs for VM-2:

Send transcript + call context

Receive response text / intent / action

API requirements:

HTTPS

Token-based authentication

Low-latency response handling

Provide API documentation:

Endpoints

Payload structure

Timeout & retry logic
Fine-Tuning & Optimization (MANDATORY)
A. Telephony & Audio Fine-Tuning

Optimize:

SIP timers

RTP jitter buffers

Packet loss handling

Eliminate:

Audio clipping

Echo

One-way audio

Tune Opus bitrate & fallback behavior

B. STT Accuracy Fine-Tuning

Tune:

Sample rate (16 kHz mono preferred)

Chunk size

Silence thresholds

Reduce:

False silence detection

Missed or cut words

Validate with real call recordings

C. TTS Naturalness Fine-Tuning

Optimize:

Pause timing

Natural speech gaps

Voice clarity over phone lines

Ensure human-like response timing

D. End-to-End Latency Optimization

Optimize full loop:

Caller → STT → AI Brain → TTS → Playback

Remove bottlenecks in:

Audio pipeline

Network calls

API responses

Goal: near-human conversational delay

E. Failure & Fallback Handling

Graceful handling for:

AI Brain timeout

STT/TTS failure

Play fallback audio instead of dropping calls

Ensure VM-1 never crashes due to external API delays

Real-World Call Testing (REQUIRED)

Freelancer must test & fine-tune for:

Silent caller

Continuous speech

Caller interrupting AI

Background noise

Long calls (10–15 minutes)

Multiple back-to-back calls

Deliverables

Fully working VM-1

SIP trunk tested with live calls

End-to-end demo call

API documentation

Fine-tuning parameters & recommendations

Deployment & restart guide

Source/config files ownership transferred

Acceptance Criteria

Calls connect reliably

Audio is clear & natural

No noticeable lag in AI responses

STT accuracy acceptable for Indian callers

TTS stops instantly on barge-in

System remains stable under repeated calls

Required Skills

SIP / RTP / VoIP

FreeSWITCH / Kamailio / OpenSIPS

Real-time audio streaming

STT & TTS systems

Linux server hardening

API design

AI voice bot experience (strong plus)

Out of Scope

Frontend or UI

LLM prompt engineering

Vector DB setup

AI Brain logic (VM-2)

Proposal Must Include

Proposed telephony stack and reason

Similar AI voice/IVR projects done

Estimated timeline

Assumptions & risks

Estimated Timeline

Setup & integration: 7–10 days

Fine-tuning & testing: 5–7 days