ASR/TTS Enhancement Expert Required (AI VOICE ENGINEER)
Budget: $15 – $25 USD
ASR/TTS Specialist (Speech In/Out)
Quick checklist
Low-latency ASR (streaming partials, solid endpointing/VAD, barge-in)
Natural TTS (fast first audio, stable prosody; SSML or equivalent controls)
Multilingual basics (accent robustness, noise handling)
Safe audio levels (no clipping; comfortable loudness)
ASR/TTS Specialist (Speed & Clarity)
Scope (weekly): Improve speech-to-text accuracy & latency; polish text-to-speech naturalness; tune streaming settings (endpointing/VAD, partials cadence, barge-in).
Must-have: Hands-on with modern streaming ASR/TTS (e.g., Deepgram/Whisper/Google/MS/AWS), SSML/prosody controls, noise/accent handling, and simple latency measurement.
Deliverables: Short weekly note with before/after metrics (p50/p95 latency, WER on a tiny test set, first-audio time for TTS) plus config diffs.
Paid test (2–4 hrs): Cut ASR p50 partial to ≤300 ms and TTS first-audio to ≤200 ms on our sample pipeline; share a tiny WER and listening sample.
Apply with: Hourly rate, 2 real voice projects (links or clips), and one paragraph on how you tune endpointing & barge-in without false cuts.
Note: You’ll pair lightly with our engineer (Alex) who handles action connectors—you focus purely on speech quality & speed.
Copy-Paste Ad — FIXED-BID
HIRING (FIXED-BID): Voice Tuning Pack (ASR/TTS)
Output:
ASR: Streaming tuned for fast partials, solid endpointing/VAD, barge-in that cuts TTS cleanly.
TTS: Fast first audio, natural prosody (SSML/equivalent), safe loudness.
Report: Before/after latency (p50/p95), small WER sample, and settings (exact params).
Handover: README + config files; “how to replicate” steps.
Milestones:
Baseline capture → record current metrics & configs
ASR tuning → partials cadence, endpointing/VAD, barge-in
TTS tuning → first-audio, prosody/SSML, loudness safety
Validation & handover → report, configs, demo clips
Acceptance (simple):
ASR partial p50 ≤300 ms, endpointing stable (no early chops); tiny test WER improved vs. baseline
TTS first-audio ≤200 ms; prosody stable; no clipping, comfortable LUFS
Barge-in reliably interrupts TTS on user speech (≥95% success in test set)
Submission: Fixed fee, short schedule, prior voice links/clips, and warranty for two quick follow-up tweaks.
Quick checklist
Low-latency ASR (streaming partials, solid endpointing/VAD, barge-in)
Natural TTS (fast first audio, stable prosody; SSML or equivalent controls)
Multilingual basics (accent robustness, noise handling)
Safe audio levels (no clipping; comfortable loudness)
ASR/TTS Specialist (Speed & Clarity)
Scope (weekly): Improve speech-to-text accuracy & latency; polish text-to-speech naturalness; tune streaming settings (endpointing/VAD, partials cadence, barge-in).
Must-have: Hands-on with modern streaming ASR/TTS (e.g., Deepgram/Whisper/Google/MS/AWS), SSML/prosody controls, noise/accent handling, and simple latency measurement.
Deliverables: Short weekly note with before/after metrics (p50/p95 latency, WER on a tiny test set, first-audio time for TTS) plus config diffs.
Paid test (2–4 hrs): Cut ASR p50 partial to ≤300 ms and TTS first-audio to ≤200 ms on our sample pipeline; share a tiny WER and listening sample.
Apply with: Hourly rate, 2 real voice projects (links or clips), and one paragraph on how you tune endpointing & barge-in without false cuts.
Note: You’ll pair lightly with our engineer (Alex) who handles action connectors—you focus purely on speech quality & speed.
Copy-Paste Ad — FIXED-BID
HIRING (FIXED-BID): Voice Tuning Pack (ASR/TTS)
Output:
ASR: Streaming tuned for fast partials, solid endpointing/VAD, barge-in that cuts TTS cleanly.
TTS: Fast first audio, natural prosody (SSML/equivalent), safe loudness.
Report: Before/after latency (p50/p95), small WER sample, and settings (exact params).
Handover: README + config files; “how to replicate” steps.
Milestones:
Baseline capture → record current metrics & configs
ASR tuning → partials cadence, endpointing/VAD, barge-in
TTS tuning → first-audio, prosody/SSML, loudness safety
Validation & handover → report, configs, demo clips
Acceptance (simple):
ASR partial p50 ≤300 ms, endpointing stable (no early chops); tiny test WER improved vs. baseline
TTS first-audio ≤200 ms; prosody stable; no clipping, comfortable LUFS
Barge-in reliably interrupts TTS on user speech (≥95% success in test set)
Submission: Fixed fee, short schedule, prior voice links/clips, and warranty for two quick follow-up tweaks.