Speaker identification / Speaker Diarization / Voice Recognition with deep learning
Budget: $750 – $1,500 USD
i need a deep learning model can reconize who spoke when with maksimum 3 speakers in same audio file.
The model can segment the speakers in same audio file by time range.
For example in the attached sample_ 1 file there are three speakers, the result should be somthing like:
A speaker: 00:00 - 00:02
B speaker: 00:02 - 00:04
A speaker: 00:05 - 00:07
B speaker: 00:07 - 00:18
A speaker: 00:18 - 00:20
B speaker: 00:20 - 00:22
A speaker: 00:22 - 00:23
B speaker: 00:23 - 00:31
A speaker: 00:31 - 00:32
B speaker: 00:32 - 00:39
C speaker: 00:39 - 01:14
B speaker: 01:14 - 01:16
C speaker: 01:16 - 01:18 .....
condetions:
* it should segment new speakers that are not in the dataset.
* in the sample_1 file the accuracy shuld be at minumum %75-80, in the sample_2 can be %65-70 and add 7575 in your proposal text.
* you can use any open source libraries or APİ even its paid but the model should has a unique part.
* I dont have dataset but there is many matarilas i will share.
* Please do your own research on the problem first.
* A successful model is expected to provide a report of at least 3 pages summarizing Materials and Methods are used.
Many Thanks
The model can segment the speakers in same audio file by time range.
For example in the attached sample_ 1 file there are three speakers, the result should be somthing like:
A speaker: 00:00 - 00:02
B speaker: 00:02 - 00:04
A speaker: 00:05 - 00:07
B speaker: 00:07 - 00:18
A speaker: 00:18 - 00:20
B speaker: 00:20 - 00:22
A speaker: 00:22 - 00:23
B speaker: 00:23 - 00:31
A speaker: 00:31 - 00:32
B speaker: 00:32 - 00:39
C speaker: 00:39 - 01:14
B speaker: 01:14 - 01:16
C speaker: 01:16 - 01:18 .....
condetions:
* it should segment new speakers that are not in the dataset.
* in the sample_1 file the accuracy shuld be at minumum %75-80, in the sample_2 can be %65-70 and add 7575 in your proposal text.
* you can use any open source libraries or APİ even its paid but the model should has a unique part.
* I dont have dataset but there is many matarilas i will share.
* Please do your own research on the problem first.
* A successful model is expected to provide a report of at least 3 pages summarizing Materials and Methods are used.
Many Thanks