Expert ASR System Developer
Budget: $10 – $30 USD
I have a production-level speech and voice recognition project that needs an experienced, hands-on engineer rather than a full agency team. The end goal is an English-only (for now) system that delivers 95–99 % word-level accuracy in real-world conditions and runs reliably during Pacific Time working hours so we can iterate together in near-real time.
What matters most is proven expertise in building complete ASR pipelines—from data collection and labeling through acoustic-language modeling, decoding, evaluation, and deployment. If your portfolio already shows models or demos powered by frameworks such as Kaldi, PyTorch, TensorFlow, or wav2vec and you can speak to the trade-offs between on-prem, AWS, GCP, or edge-level inference, that experience will fit perfectly.
Because this is an independent-contractor engagement, I’ll be collaborating with you directly on model architecture, domain adaptation, and error analysis. Clear version control (Git), reproducible training scripts, and concise experiment tracking are non-negotiable so we can benchmark progress against the 95 %+ target. Should future phases require bilingual or multilingual expansion, the groundwork you lay now should make adding Spanish or other languages straightforward.
For an initial milestone, I’d like to see:
• a brief technical plan outlining data requirements, preprocessing approach, and baseline model choice
• an early prototype trained on a public English corpus, accompanied by a WER report that proves we’re on track for the accuracy target
• a short video or live demo of the recognizer handling both command-style utterances and continuous speech
Please share links or repositories that illustrate previous ASR or speech-driven projects so we can hit the ground running. I’m based in Los Angeles (UTC-8) and will expect regular syncs within that window. Looking forward to seeing how your past work can translate into a robust recognizer for this new build.
What matters most is proven expertise in building complete ASR pipelines—from data collection and labeling through acoustic-language modeling, decoding, evaluation, and deployment. If your portfolio already shows models or demos powered by frameworks such as Kaldi, PyTorch, TensorFlow, or wav2vec and you can speak to the trade-offs between on-prem, AWS, GCP, or edge-level inference, that experience will fit perfectly.
Because this is an independent-contractor engagement, I’ll be collaborating with you directly on model architecture, domain adaptation, and error analysis. Clear version control (Git), reproducible training scripts, and concise experiment tracking are non-negotiable so we can benchmark progress against the 95 %+ target. Should future phases require bilingual or multilingual expansion, the groundwork you lay now should make adding Spanish or other languages straightforward.
For an initial milestone, I’d like to see:
• a brief technical plan outlining data requirements, preprocessing approach, and baseline model choice
• an early prototype trained on a public English corpus, accompanied by a WER report that proves we’re on track for the accuracy target
• a short video or live demo of the recognizer handling both command-style utterances and continuous speech
Please share links or repositories that illustrate previous ASR or speech-driven projects so we can hit the ground running. I’m based in Los Angeles (UTC-8) and will expect regular syncs within that window. Looking forward to seeing how your past work can translate into a robust recognizer for this new build.