Kaldi ASR Toolkit Customization
Budget: $200 – $500 USD
Project Description:
We are looking for an experienced AI engineer or sound engineer with hands-on experience working with Automatic Speech Recognition (ASR) toolkits—specifically Kaldi.
Scope of Work:
Customize and optimize Kaldi toolkit scripts (data preparation, language modeling, acoustic model training, decoding).
Handle corpora like WSJ, with .wv1 audio files and .dot/.ptx transcriptions.
Generate metadata (wav.scp, text, utt2spk, spk2utt) and configure lexicons, language models (ARPA), and decoding graphs (HCLG.fst).
Support dialectal or multi-lingual ASR setups (preferably Modern Standard Arabic and dialects).
Debug and document errors in training pipelines, feature extraction, and decoding results.
(Bonus) Help convert models for integration with real-time applications or export to ONNX/other formats.
Requirements:
Strong background in Kaldi (or equivalent ASR frameworks like ESPnet, Vosk, Whisper, etc.)
Familiarity with data structures and scripting in Bash and Python.
Experience working with large ASR corpora (WSJ, Fisher, LibriSpeech, etc.).
Familiarity with signal processing concepts and feature extraction (MFCC, FBank, etc.).
Good communication and documentation skills.
Preferred but not required:
Arabic ASR experience (MSA and dialects)
Knowledge of language modeling (SRILM or KenLM)
Ability to build and tune chain models or neural network-based acoustic models (nnet3 or TDNN)
Deliverables:
Working Kaldi pipeline for training and decoding.
Full documentation of steps and configuration files.
Support for integration or testing if needed.
Project Type:
One-time project with possibility for future collaboration.
Duration:
Expected to be completed in 2–4 weeks.
To Apply:
Please include:
Summary of your experience with Kaldi or similar toolkits.
Links to relevant GitHub repos or past ASR work.
Your proposed approach and availability.
We are looking for an experienced AI engineer or sound engineer with hands-on experience working with Automatic Speech Recognition (ASR) toolkits—specifically Kaldi.
Scope of Work:
Customize and optimize Kaldi toolkit scripts (data preparation, language modeling, acoustic model training, decoding).
Handle corpora like WSJ, with .wv1 audio files and .dot/.ptx transcriptions.
Generate metadata (wav.scp, text, utt2spk, spk2utt) and configure lexicons, language models (ARPA), and decoding graphs (HCLG.fst).
Support dialectal or multi-lingual ASR setups (preferably Modern Standard Arabic and dialects).
Debug and document errors in training pipelines, feature extraction, and decoding results.
(Bonus) Help convert models for integration with real-time applications or export to ONNX/other formats.
Requirements:
Strong background in Kaldi (or equivalent ASR frameworks like ESPnet, Vosk, Whisper, etc.)
Familiarity with data structures and scripting in Bash and Python.
Experience working with large ASR corpora (WSJ, Fisher, LibriSpeech, etc.).
Familiarity with signal processing concepts and feature extraction (MFCC, FBank, etc.).
Good communication and documentation skills.
Preferred but not required:
Arabic ASR experience (MSA and dialects)
Knowledge of language modeling (SRILM or KenLM)
Ability to build and tune chain models or neural network-based acoustic models (nnet3 or TDNN)
Deliverables:
Working Kaldi pipeline for training and decoding.
Full documentation of steps and configuration files.
Support for integration or testing if needed.
Project Type:
One-time project with possibility for future collaboration.
Duration:
Expected to be completed in 2–4 weeks.
To Apply:
Please include:
Summary of your experience with Kaldi or similar toolkits.
Links to relevant GitHub repos or past ASR work.
Your proposed approach and availability.
Related categories:
Python
Debugging
Neural Networks
Documentation
Signal Processing
Bash
Automatic Speech Recognition