Speech Datasets Needed (19 Languages, Reading Dialogues)
Budget: $30 – $250 USD
We are seeking finished, ready-to-use speech datasets in the following 19 languages/dialects, along with their specific country/region, for a client project.
Languages & Regions Needed:
Achinese – Indonesia
Acehnese (Malaysia) – Malaysia
Acehnese (Thailand) – Thailand
Balinese – Indonesia
Buginese – Indonesia
Makassarese – Indonesia
Minangkabau – Indonesia
Sasak – Indonesia
Sundanese – Indonesia
Toraja – Indonesia
Ambonese Malay – Indonesia
Javanese – Indonesia
Madurese – Indonesia
Betawi Malay – Indonesia
Banyumasan – Indonesia
Bantenese – Indonesia
Cirebonese – Indonesia
Osing – Indonesia
Tenggerese – Indonesia
Requirements:
Content type: Reading dialogues (audio recordings of scripted or natural conversations)
Format: Audio + transcription (text aligned with the audio)
Quantity: No minimum or maximum — any amount is acceptable
Quality: Clear audio, minimal background noise, native speaker accents preferred
Usage rights: Must have full rights to resell or license the dataset for commercial use
Notes:
This request is for finished datasets only — no new data collection required at this stage
No sample files are required for the initial discussion
Please indicate for each language:
Language name
Total hours available
Price per hour or total price
File format(s)
How to Apply:
Send a message with:
The languages you have available
The dataset size (in hours) for each
Your asking price
Confirmation that you have commercial rights to sell/license the data
Languages & Regions Needed:
Achinese – Indonesia
Acehnese (Malaysia) – Malaysia
Acehnese (Thailand) – Thailand
Balinese – Indonesia
Buginese – Indonesia
Makassarese – Indonesia
Minangkabau – Indonesia
Sasak – Indonesia
Sundanese – Indonesia
Toraja – Indonesia
Ambonese Malay – Indonesia
Javanese – Indonesia
Madurese – Indonesia
Betawi Malay – Indonesia
Banyumasan – Indonesia
Bantenese – Indonesia
Cirebonese – Indonesia
Osing – Indonesia
Tenggerese – Indonesia
Requirements:
Content type: Reading dialogues (audio recordings of scripted or natural conversations)
Format: Audio + transcription (text aligned with the audio)
Quantity: No minimum or maximum — any amount is acceptable
Quality: Clear audio, minimal background noise, native speaker accents preferred
Usage rights: Must have full rights to resell or license the dataset for commercial use
Notes:
This request is for finished datasets only — no new data collection required at this stage
No sample files are required for the initial discussion
Please indicate for each language:
Language name
Total hours available
Price per hour or total price
File format(s)
How to Apply:
Send a message with:
The languages you have available
The dataset size (in hours) for each
Your asking price
Confirmation that you have commercial rights to sell/license the data