Record Japanese Dialogues
Budget: ₹1,500 – ₹12,500 INR
Data Collection Specification
Mock-up call center conversation data
1. The speech data will be collected by data collectors with natural speaking by mocking-up the bank call center conversations between Client and Agent roles based on the provided script.
2. Script reading will be not acceptable. Collectors can get familiar with the script first, then understand it themselves, and then have a free conversation based on it as long as you don't stray from the topic. There should be tone words, natural pauses, overlapping speech included in the recordings. Please note, you cannot read the script as it is, otherwise we will consider it as invalid audio.
3. The data collectors should be Japanese native speakers.
4. One pair collector can have at most 60 minutes audio recordings, each all can be 10 minutes.
5. Some background noise has to be audible in the recording. The background noise can be any type of ambient sound like television playing, but it must not overpower or cover the participants' voices. The conversation content should remain clear and easy to understand.
5. Speakers gender distribution in a relatively balanced proportion of 50% Female and 50% Male, ±10% can acceptable.
6. The valid speech length from data collectors (by removing the silence and noisy) would be no less than 80% of the audio length.
7. You can record via phone call, internet call, or other similar way.
8. Recording name format:
Dialogue ID- Speaker A’s name-Speaker B’s name
for example: Japanese-1-Henry-Linda.wav
Mock-up call center conversation data
1. The speech data will be collected by data collectors with natural speaking by mocking-up the bank call center conversations between Client and Agent roles based on the provided script.
2. Script reading will be not acceptable. Collectors can get familiar with the script first, then understand it themselves, and then have a free conversation based on it as long as you don't stray from the topic. There should be tone words, natural pauses, overlapping speech included in the recordings. Please note, you cannot read the script as it is, otherwise we will consider it as invalid audio.
3. The data collectors should be Japanese native speakers.
4. One pair collector can have at most 60 minutes audio recordings, each all can be 10 minutes.
5. Some background noise has to be audible in the recording. The background noise can be any type of ambient sound like television playing, but it must not overpower or cover the participants' voices. The conversation content should remain clear and easy to understand.
5. Speakers gender distribution in a relatively balanced proportion of 50% Female and 50% Male, ±10% can acceptable.
6. The valid speech length from data collectors (by removing the silence and noisy) would be no less than 80% of the audio length.
7. You can record via phone call, internet call, or other similar way.
8. Recording name format:
Dialogue ID- Speaker A’s name-Speaker B’s name
for example: Japanese-1-Henry-Linda.wav