Pre-Recorded Call Center Conversations Dataset Needed (English, Hindi, Spanish)
Budget: $10 – $5,000 USD
We are looking for high-quality pre-recorded call center conversation datasets for AI/ASR model training purposes.
Languages Required:
• English – 500 Hours
• Hindi – 500 Hours
• Spanish – 500 Hours
Dataset Requirements:
• Call center conversation audio
• Clean audio without background noise
• No long silence segments
• WAV audio format
• 16 kHz, 16-bit or higher
• Unidirectional audio preferred
• Transcription text required
Please provide the following information with your proposal:
1. Transcription Details
• Is transcription text available?
• Is the transcription AI-generated or human-annotated?
• If AI-generated, can it be manually reviewed/refined?
2. Metadata & Accuracy Information
Please confirm whether the dataset includes:
• Speaker tags
• Timestamp information
• Speaker diarization details
Also share accuracy metrics for:
• Speaker diarization
• WER (Word Error Rate)
• Timestamp alignment
3. Audio Specifications
• Confirm sampling rate and audio quality details.
4. Data Privacy & Compliance
• Has all personal or sensitive information been de-identified/anonymized?
• Please explain the de-identification process used.
5. Usage Scope
• The dataset will be used strictly for internal AI/model training purposes only.
Important Notes:
• We are only looking for pre-recorded/off-the-shelf datasets.
• Fresh recordings or newly collected data are not required for this project.
• Please share sample files, pricing, delivery timeline, and licensing/commercial usage terms if available.
If you have relevant datasets available, feel free to send details via message or proposal.
Languages Required:
• English – 500 Hours
• Hindi – 500 Hours
• Spanish – 500 Hours
Dataset Requirements:
• Call center conversation audio
• Clean audio without background noise
• No long silence segments
• WAV audio format
• 16 kHz, 16-bit or higher
• Unidirectional audio preferred
• Transcription text required
Please provide the following information with your proposal:
1. Transcription Details
• Is transcription text available?
• Is the transcription AI-generated or human-annotated?
• If AI-generated, can it be manually reviewed/refined?
2. Metadata & Accuracy Information
Please confirm whether the dataset includes:
• Speaker tags
• Timestamp information
• Speaker diarization details
Also share accuracy metrics for:
• Speaker diarization
• WER (Word Error Rate)
• Timestamp alignment
3. Audio Specifications
• Confirm sampling rate and audio quality details.
4. Data Privacy & Compliance
• Has all personal or sensitive information been de-identified/anonymized?
• Please explain the de-identification process used.
5. Usage Scope
• The dataset will be used strictly for internal AI/model training purposes only.
Important Notes:
• We are only looking for pre-recorded/off-the-shelf datasets.
• Fresh recordings or newly collected data are not required for this project.
• Please share sample files, pricing, delivery timeline, and licensing/commercial usage terms if available.
If you have relevant datasets available, feel free to send details via message or proposal.