ASR/TTS Enhancement Expert Required (AI VOICE ENGINEER)

Job ID: 39919997

Budget: $15 – $25 USD

ASR/TTS Specialist (Speech In/Out)




Quick checklist

Low-latency ASR (streaming partials, solid endpointing/VAD, barge-in)

Natural TTS (fast first audio, stable prosody; SSML or equivalent controls)

Multilingual basics (accent robustness, noise handling)

Safe audio levels (no clipping; comfortable loudness)

ASR/TTS Specialist (Speed & Clarity)

Scope (weekly): Improve speech-to-text accuracy & latency; polish text-to-speech naturalness; tune streaming settings (endpointing/VAD, partials cadence, barge-in).
Must-have: Hands-on with modern streaming ASR/TTS (e.g., Deepgram/Whisper/Google/MS/AWS), SSML/prosody controls, noise/accent handling, and simple latency measurement.
Deliverables: Short weekly note with before/after metrics (p50/p95 latency, WER on a tiny test set, first-audio time for TTS) plus config diffs.
Paid test (2–4 hrs): Cut ASR p50 partial to ≤300 ms and TTS first-audio to ≤200 ms on our sample pipeline; share a tiny WER and listening sample.
Apply with: Hourly rate, 2 real voice projects (links or clips), and one paragraph on how you tune endpointing & barge-in without false cuts.
Note: You’ll pair lightly with our engineer (Alex) who handles action connectors—you focus purely on speech quality & speed.

Copy-Paste Ad — FIXED-BID

HIRING (FIXED-BID): Voice Tuning Pack (ASR/TTS)

Output:

ASR: Streaming tuned for fast partials, solid endpointing/VAD, barge-in that cuts TTS cleanly.

TTS: Fast first audio, natural prosody (SSML/equivalent), safe loudness.

Report: Before/after latency (p50/p95), small WER sample, and settings (exact params).

Handover: README + config files; “how to replicate” steps.

Milestones:

Baseline capture → record current metrics & configs

ASR tuning → partials cadence, endpointing/VAD, barge-in

TTS tuning → first-audio, prosody/SSML, loudness safety

Validation & handover → report, configs, demo clips

Acceptance (simple):

ASR partial p50 ≤300 ms, endpointing stable (no early chops); tiny test WER improved vs. baseline

TTS first-audio ≤200 ms; prosody stable; no clipping, comfortable LUFS

Barge-in reliably interrupts TTS on user speech (≥95% success in test set)

Submission: Fixed fee, short schedule, prior voice links/clips, and warranty for two quick follow-up tweaks.