Call Center Audio Datasets Needed English , Hindi , Spanish
Budget: ₹750 – ₹1,250 INR
We are currently looking for pre-recorded call center conversation datasets with the following requirements:
• English – 500 hours
• Hindi – 500 hours
• Spanish – 500 hours
Requirements:
* Agent-side audio only
* No background noise
* No long silence segments
* Audio format: WAV
* 16 kHz, 16-bit or higher
* Unidirectional audio
* Transcription text required
Additionally, please confirm the following details:
1. Is transcription text available?
* If yes, is it AI-generated or human-annotated?
* If AI-generated, can it be manually refined?
2. If transcripts are available, please share details regarding:
* Speaker tags
* Timestamp information
* Accuracy rates for:
• Speaker diarization
• WER (Word Error Rate)
• Timestamp alignment
3. Please confirm the audio sampling rate specifications.
4. Has personal information within the recordings been de-identified?
* If yes, please explain the de-identification process.
5. Scope of Data Use:
* The dataset will only be used for internal model training purposes.
• English – 500 hours
• Hindi – 500 hours
• Spanish – 500 hours
Requirements:
* Agent-side audio only
* No background noise
* No long silence segments
* Audio format: WAV
* 16 kHz, 16-bit or higher
* Unidirectional audio
* Transcription text required
Additionally, please confirm the following details:
1. Is transcription text available?
* If yes, is it AI-generated or human-annotated?
* If AI-generated, can it be manually refined?
2. If transcripts are available, please share details regarding:
* Speaker tags
* Timestamp information
* Accuracy rates for:
• Speaker diarization
• WER (Word Error Rate)
• Timestamp alignment
3. Please confirm the audio sampling rate specifications.
4. Has personal information within the recordings been de-identified?
* If yes, please explain the de-identification process.
5. Scope of Data Use:
* The dataset will only be used for internal model training purposes.