Speech Datasets Needed (19 Languages, Reading Dialogues)

Job ID: 39699047

Budget: $30 – $250 USD

We are seeking finished, ready-to-use speech datasets in the following 19 languages/dialects, along with their specific country/region, for a client project.

Languages & Regions Needed:

Achinese – Indonesia

Acehnese (Malaysia) – Malaysia

Acehnese (Thailand) – Thailand

Balinese – Indonesia

Buginese – Indonesia

Makassarese – Indonesia

Minangkabau – Indonesia

Sasak – Indonesia

Sundanese – Indonesia

Toraja – Indonesia

Ambonese Malay – Indonesia

Javanese – Indonesia

Madurese – Indonesia

Betawi Malay – Indonesia

Banyumasan – Indonesia

Bantenese – Indonesia

Cirebonese – Indonesia

Osing – Indonesia

Tenggerese – Indonesia

Requirements:

Content type: Reading dialogues (audio recordings of scripted or natural conversations)

Format: Audio + transcription (text aligned with the audio)

Quantity: No minimum or maximum — any amount is acceptable

Quality: Clear audio, minimal background noise, native speaker accents preferred

Usage rights: Must have full rights to resell or license the dataset for commercial use

Notes:

This request is for finished datasets only — no new data collection required at this stage

No sample files are required for the initial discussion

Please indicate for each language:

Language name

Total hours available

Price per hour or total price

File format(s)

How to Apply:
Send a message with:

The languages you have available

The dataset size (in hours) for each

Your asking price

Confirmation that you have commercial rights to sell/license the data