Python Dev Needed: Real-Time Audio/Selenium Fixes + n8n Integration (5-Hour Deadline)
Budget: €150 – €200 EUR
The project is a self-hosted bot that automatically joins online meetings (Teams → Meet → Zoom), waits for a wake-word, records the spoken sentence, routes it to ElevenLabs STT → n8n (GPT + DB) and speaks the generated reply back into the call through a virtual microphone. All steps must happen with the lowest possible latency (< 2 s ideal).
The codebase already contains:
Selenium bot for Microsoft Teams
Wake-word detector (Picovoice Porcupine)
ElevenLabs STT integration
n8n workflow that returns a ready-to-play MP3
We need a freelancer to finish the missing pieces, harden what is there and push it to production.
Deliverables
Reliable audio I/O
Capture mixed meeting audio in real time (Teams / Meet / Zoom).
Play back MP3 responses into the call through the virtual microphone.
Verify end-to-end latency stays below 2 seconds.
Additional platform support
Implement the headless-browser bot for Google Meet and Zoom (Teams already works).
Auto-join via invite link (sent in the client).
Interruption handling
When the bot is speaking and someone utters the wake-word, immediately stop playback for that phrase.
Latency optimisation
Profile each stage (wake-word → STT → n8n → TTS → playback) and trim overhead.
Aim for < 1.5 s round-trip in ideal conditions.
Timeline
All deliverables must be completed within a maximum of 5 hours from contract start.
Must-have skills
Strong Python (asyncio, multithreading, websockets)
Selenium / Playwright or similar headless-browser automation
Deep understanding of real-time audio on Linux (PulseAudio / PipeWire)
Experience with WebRTC, VAD and basic DSP
Docker
Comfortable reading n8n workflows / REST
Access & Assets Provided
All necessary API keys will be provided by us.
You will receive a ZIP archive with two folders:
n8n – the workflow and related assets
python – client and server code (kept in separate subfolders)
The codebase already contains:
Selenium bot for Microsoft Teams
Wake-word detector (Picovoice Porcupine)
ElevenLabs STT integration
n8n workflow that returns a ready-to-play MP3
We need a freelancer to finish the missing pieces, harden what is there and push it to production.
Deliverables
Reliable audio I/O
Capture mixed meeting audio in real time (Teams / Meet / Zoom).
Play back MP3 responses into the call through the virtual microphone.
Verify end-to-end latency stays below 2 seconds.
Additional platform support
Implement the headless-browser bot for Google Meet and Zoom (Teams already works).
Auto-join via invite link (sent in the client).
Interruption handling
When the bot is speaking and someone utters the wake-word, immediately stop playback for that phrase.
Latency optimisation
Profile each stage (wake-word → STT → n8n → TTS → playback) and trim overhead.
Aim for < 1.5 s round-trip in ideal conditions.
Timeline
All deliverables must be completed within a maximum of 5 hours from contract start.
Must-have skills
Strong Python (asyncio, multithreading, websockets)
Selenium / Playwright or similar headless-browser automation
Deep understanding of real-time audio on Linux (PulseAudio / PipeWire)
Experience with WebRTC, VAD and basic DSP
Docker
Comfortable reading n8n workflows / REST
Access & Assets Provided
All necessary API keys will be provided by us.
You will receive a ZIP archive with two folders:
n8n – the workflow and related assets
python – client and server code (kept in separate subfolders)