Pre-Recorded Call Center Conversations Dataset Needed (English, Hindi, Spanish)

Job ID: 40475844

Budget: $10 – $5,000 USD

We are looking for high-quality pre-recorded call center conversation datasets for AI/ASR model training purposes.

Languages Required:
• English – 500 Hours
• Hindi – 500 Hours
• Spanish – 500 Hours

Dataset Requirements:
• Call center conversation audio
• Clean audio without background noise
• No long silence segments
• WAV audio format
• 16 kHz, 16-bit or higher
• Unidirectional audio preferred
• Transcription text required

Please provide the following information with your proposal:

1. Transcription Details
• Is transcription text available?
• Is the transcription AI-generated or human-annotated?
• If AI-generated, can it be manually reviewed/refined?

2. Metadata & Accuracy Information
Please confirm whether the dataset includes:
• Speaker tags
• Timestamp information
• Speaker diarization details

Also share accuracy metrics for:
• Speaker diarization
• WER (Word Error Rate)
• Timestamp alignment

3. Audio Specifications
• Confirm sampling rate and audio quality details.

4. Data Privacy & Compliance
• Has all personal or sensitive information been de-identified/anonymized?
• Please explain the de-identification process used.

5. Usage Scope
• The dataset will be used strictly for internal AI/model training purposes only.

Important Notes:
• We are only looking for pre-recorded/off-the-shelf datasets.
• Fresh recordings or newly collected data are not required for this project.
• Please share sample files, pricing, delivery timeline, and licensing/commercial usage terms if available.

If you have relevant datasets available, feel free to send details via message or proposal.