Real-Time Speaker Diarization & Transcription AI

Job ID: 38631975

Budget: $30 – $250 USD

I am looking for a developer to build an AI agent that can perform real-time speaker diarization and transcription using Whisper for transcription. The agent should be able to distinguish between multiple speakers (e.g., Speaker A, Speaker B, Speaker C) in real time, not by uploading a pre-recorded file but during live speech.

Key Features:
- Real-time processing of live audio input
- Basic separation of voice speakers, not needing named identification or integration with user profiles

Ideal Skills:
- Experience with AI and machine learning
- Proficiency in audio processing and transcription technologies
- Familiarity with Whisper

The goal is to have a robust system for basic speaker differentiation in a controlled environment - a conference room.