Voice Gateway & Telephony Integration for Rasa AI Assistant
Budget: ₹12,500 – ₹37,500 INR
We have a fully trained Rasa 3.x conversational model and now want to build the Voice Gateway & Telephony layer that enables:
• Real-time speech-to-text (ASR) and text-to-speech (TTS)
• WebRTC/WebSocket audio streaming from browser
• Phone call integration (Twilio Media Streams) with barge-in and low latency
• Seamless connection with the Rasa bot (Socket.IO/REST)
This is Module 1 of a multi-phase project. Successful completion can lead to further modules (Web UI, analytics, autoscaling, etc.).
⸻
Responsibilities
1. Voice Gateway service
• Develop a backend (Node.js/TS or Python/FastAPI) that:
• Receives audio from browser (WebRTC) and phone calls (Twilio Media Streams)
• Converts speech→text (ASR) and text→speech (TTS) in real-time
• Handles barge-in (interrupt bot) and endpointing (VAD)
• Communicates with the Rasa server via Socket.IO or REST
2. Telephony
• Set up a Twilio number and Media Streams connection to the Voice Gateway
• Manage call session lifecycle, retries, and error handling
3. Performance
• Target latency: ≤2s time-to-first-audio
• Barge-in response: ≤250 ms stop when user interrupts
4. Documentation & Handover
• Dockerized service + clear runbook
• Architecture diagram, setup guide, and code walkthrough
⸻
Requirements
• Strong experience with real-time audio streaming (WebRTC or Twilio Media Streams)
• Prior work with ASR/TTS integrations (Google/Azure/Deepgram/ElevenLabs or self-hosted faster-whisper/Piper)
• Familiarity with Rasa or any chatbot orchestration framework (preferred)
• Good knowledge of Node.js/TypeScript or Python (FastAPI)
• Experience with Docker, WebSockets, and scalable backend design
• Comfortable working in IST time zone with 2 hours overlap with ET
⸻
Deliverables for this Module
• Voice Gateway backend that supports web audio and phone calls
• ASR + TTS pipeline with streaming and barge-in
• Working end-to-end demo:
• User can talk via browser → bot responds via voice
• User can call Twilio number → bot responds via voice
• Dockerized deployment with basic logging/metrics
• Handover documentation
Bonus if you have:
• Experience with barge-in / VAD tuning
• Built voice bots integrated with Twilio before
• Experience with Dockerized microservices and basic observability
• Real-time speech-to-text (ASR) and text-to-speech (TTS)
• WebRTC/WebSocket audio streaming from browser
• Phone call integration (Twilio Media Streams) with barge-in and low latency
• Seamless connection with the Rasa bot (Socket.IO/REST)
This is Module 1 of a multi-phase project. Successful completion can lead to further modules (Web UI, analytics, autoscaling, etc.).
⸻
Responsibilities
1. Voice Gateway service
• Develop a backend (Node.js/TS or Python/FastAPI) that:
• Receives audio from browser (WebRTC) and phone calls (Twilio Media Streams)
• Converts speech→text (ASR) and text→speech (TTS) in real-time
• Handles barge-in (interrupt bot) and endpointing (VAD)
• Communicates with the Rasa server via Socket.IO or REST
2. Telephony
• Set up a Twilio number and Media Streams connection to the Voice Gateway
• Manage call session lifecycle, retries, and error handling
3. Performance
• Target latency: ≤2s time-to-first-audio
• Barge-in response: ≤250 ms stop when user interrupts
4. Documentation & Handover
• Dockerized service + clear runbook
• Architecture diagram, setup guide, and code walkthrough
⸻
Requirements
• Strong experience with real-time audio streaming (WebRTC or Twilio Media Streams)
• Prior work with ASR/TTS integrations (Google/Azure/Deepgram/ElevenLabs or self-hosted faster-whisper/Piper)
• Familiarity with Rasa or any chatbot orchestration framework (preferred)
• Good knowledge of Node.js/TypeScript or Python (FastAPI)
• Experience with Docker, WebSockets, and scalable backend design
• Comfortable working in IST time zone with 2 hours overlap with ET
⸻
Deliverables for this Module
• Voice Gateway backend that supports web audio and phone calls
• ASR + TTS pipeline with streaming and barge-in
• Working end-to-end demo:
• User can talk via browser → bot responds via voice
• User can call Twilio number → bot responds via voice
• Dockerized deployment with basic logging/metrics
• Handover documentation
Bonus if you have:
• Experience with barge-in / VAD tuning
• Built voice bots integrated with Twilio before
• Experience with Dockerized microservices and basic observability
Related categories:
Python
Asterisk PBX
VoIP
Node.js
XMPP
Docker
Twilio
Microservices
WebRTC
FastAPI