Call Center Speech Dataset Collection (English, Hindi, Spanish )

Job ID: 40470137

Budget: $10 – $15,000 USD

Project Description

We are looking for experienced freelancers or teams to provide or build a high-quality call center conversation dataset for AI training.

We are open to both:

* Custom data collection and annotation
* Ready-made (pre-built) datasets that meet requirements

-Languages Required

* English – 500 hours
* Hindi – 500 hours
* Spanish – 500 hours

Total: 1500 hours

Data Requirements

* Call center style agent–customer conversations only
* Clean audio (no background noise)
* Minimal silence and natural flow
* WAV format
* 16 kHz sampling rate (16-bit or higher)
* Single channel preferred

Transcription Requirements

* Full transcription required
* Speaker labels (Agent / Customer)
* Timestamped alignment required

Metadata Requirements (Must Confirm)

* Source type (recorded or pre-built dataset)
* Transcription type (AI or human-annotated)
* Ability to refine AI transcripts if applicable
* Speaker diarization accuracy
* Timestamp alignment accuracy
* Estimated WER (Word Error Rate)
* Audio sampling rate details

Privacy & Compliance

* All personal data must be anonymized or removed
* Method of de-identification must be clearly explained
* Data must be legally collected with proper consent
* Only for internal AI model training use

Deliverables

* WAV audio files
* Transcript files (TXT or JSON)
* Metadata (speaker labels, timestamps, language info)

Requirements from Freelancer

Please include:

* Experience with speech datasets or ASR projects
* Whether you can provide ready-made datasets
* Tools and workflow used
* Production capacity
* Estimated cost per hour
* Delivery timeline

Important Note

We are open to **ready-made datasets as long as they fully meet the above requirements and are legally authorized for AI training use**.