Real-Time Speaker Diarization & Transcription AI
Budget: $30 – $250 USD
I am looking for a developer to build an AI agent that can perform real-time speaker diarization and transcription using Whisper for transcription. The agent should be able to distinguish between multiple speakers (e.g., Speaker A, Speaker B, Speaker C) in real time, not by uploading a pre-recorded file but during live speech.
Key Features:
- Real-time processing of live audio input
- Basic separation of voice speakers, not needing named identification or integration with user profiles
Ideal Skills:
- Experience with AI and machine learning
- Proficiency in audio processing and transcription technologies
- Familiarity with Whisper
The goal is to have a robust system for basic speaker differentiation in a controlled environment - a conference room.
Key Features:
- Real-time processing of live audio input
- Basic separation of voice speakers, not needing named identification or integration with user profiles
Ideal Skills:
- Experience with AI and machine learning
- Proficiency in audio processing and transcription technologies
- Familiarity with Whisper
The goal is to have a robust system for basic speaker differentiation in a controlled environment - a conference room.