Call Center Speech Dataset Collection (English, Hindi, Spanish )
Budget: $10 – $15,000 USD
Project Description
We are looking for experienced freelancers or teams to provide or build a high-quality call center conversation dataset for AI training.
We are open to both:
* Custom data collection and annotation
* Ready-made (pre-built) datasets that meet requirements
-Languages Required
* English – 500 hours
* Hindi – 500 hours
* Spanish – 500 hours
Total: 1500 hours
Data Requirements
* Call center style agent–customer conversations only
* Clean audio (no background noise)
* Minimal silence and natural flow
* WAV format
* 16 kHz sampling rate (16-bit or higher)
* Single channel preferred
Transcription Requirements
* Full transcription required
* Speaker labels (Agent / Customer)
* Timestamped alignment required
Metadata Requirements (Must Confirm)
* Source type (recorded or pre-built dataset)
* Transcription type (AI or human-annotated)
* Ability to refine AI transcripts if applicable
* Speaker diarization accuracy
* Timestamp alignment accuracy
* Estimated WER (Word Error Rate)
* Audio sampling rate details
Privacy & Compliance
* All personal data must be anonymized or removed
* Method of de-identification must be clearly explained
* Data must be legally collected with proper consent
* Only for internal AI model training use
Deliverables
* WAV audio files
* Transcript files (TXT or JSON)
* Metadata (speaker labels, timestamps, language info)
Requirements from Freelancer
Please include:
* Experience with speech datasets or ASR projects
* Whether you can provide ready-made datasets
* Tools and workflow used
* Production capacity
* Estimated cost per hour
* Delivery timeline
Important Note
We are open to **ready-made datasets as long as they fully meet the above requirements and are legally authorized for AI training use**.
We are looking for experienced freelancers or teams to provide or build a high-quality call center conversation dataset for AI training.
We are open to both:
* Custom data collection and annotation
* Ready-made (pre-built) datasets that meet requirements
-Languages Required
* English – 500 hours
* Hindi – 500 hours
* Spanish – 500 hours
Total: 1500 hours
Data Requirements
* Call center style agent–customer conversations only
* Clean audio (no background noise)
* Minimal silence and natural flow
* WAV format
* 16 kHz sampling rate (16-bit or higher)
* Single channel preferred
Transcription Requirements
* Full transcription required
* Speaker labels (Agent / Customer)
* Timestamped alignment required
Metadata Requirements (Must Confirm)
* Source type (recorded or pre-built dataset)
* Transcription type (AI or human-annotated)
* Ability to refine AI transcripts if applicable
* Speaker diarization accuracy
* Timestamp alignment accuracy
* Estimated WER (Word Error Rate)
* Audio sampling rate details
Privacy & Compliance
* All personal data must be anonymized or removed
* Method of de-identification must be clearly explained
* Data must be legally collected with proper consent
* Only for internal AI model training use
Deliverables
* WAV audio files
* Transcript files (TXT or JSON)
* Metadata (speaker labels, timestamps, language info)
Requirements from Freelancer
Please include:
* Experience with speech datasets or ASR projects
* Whether you can provide ready-made datasets
* Tools and workflow used
* Production capacity
* Estimated cost per hour
* Delivery timeline
Important Note
We are open to **ready-made datasets as long as they fully meet the above requirements and are legally authorized for AI training use**.