Decipher Multi-Speaker Audio Recordings

Job ID: 40079611

Budget: $30 – $250 USD

I have an audio file that captures a lively exchange involving more than two people. My priority is to see every spoken word converted into clear, well-formatted text, with each speaker reliably tagged so the conversation reads naturally and is easy to follow.

Alongside the dialogue, please call out any other audible human sounds you hear—laughter, sighs, throat-clearing, even an indistinct shout in the background—so I can understand the full context of the recording. Non-human ambient noise can be ignored unless it interferes with deciphering the speech.

Deliverable
• A time-stamped transcript in .docx or .txt
• Distinct speaker labels (Speaker 1, Speaker 2, etc.) for all voices
• Bracketed notes for notable human sounds: [laughter], [cough], [unclear male voice in background], and similar

Acceptance criteria
• 98 %+ word-level accuracy for the spoken conversation
• Consistent speaker differentiation throughout
• All human sounds captured or explicitly marked as inaudible

If you work with tools like Express Scribe, Audacity or similar software that helps identify overlapping voices, feel free to use them—accuracy is what matters most to me. Turnaround within a few days is ideal, but let me know your realistic timeframe when you bid.