AI Voice Assistant for Computers -- 4
Budget: $3,000 – $5,000 USD
I’m building a voice-driven virtual assistant that runs natively on desktop and laptop machines (Windows, macOS, and ideally Linux). The core idea is simple: once installed, the user should be able to say a wake word, speak naturally, and have the agent interpret the request, fetch or process the needed information, then respond aloud and/or execute system-level actions (opening apps, searching the web, reading email summaries, controlling media, etc.).
Key needs
• Accurate, low-latency speech-to-text and natural-language understanding
• Text-to-speech replies that sound natural and can switch between at least two voices
• An intent framework that I can extend with new skills via simple Python modules or REST hooks
• Local caching of personal data and settings so the assistant works offline for basic tasks
• A streamlined installer and a lightweight tray/menu-bar interface for settings
Deliverables I’ll review for acceptance:
- A compiled desktop application and accompanying source code
- Clear build/run instructions plus a short user guide
- A brief demo video showing the wake word, two different commands, and the assistant’s spoken responses
I’m comfortable with you leveraging open-source libraries such as Vosk, Whisper, Rasa, or similar; just document what you choose and why. Clean architecture, readable code, and future extensibility matter more to me than flashy UI. If you’ve shipped anything comparable—desktop dictation tools, Jarvis-style assistants, or speech-enabled productivity apps—let me know, and point me to a live demo or repo.
Key needs
• Accurate, low-latency speech-to-text and natural-language understanding
• Text-to-speech replies that sound natural and can switch between at least two voices
• An intent framework that I can extend with new skills via simple Python modules or REST hooks
• Local caching of personal data and settings so the assistant works offline for basic tasks
• A streamlined installer and a lightweight tray/menu-bar interface for settings
Deliverables I’ll review for acceptance:
- A compiled desktop application and accompanying source code
- Clear build/run instructions plus a short user guide
- A brief demo video showing the wake word, two different commands, and the assistant’s spoken responses
I’m comfortable with you leveraging open-source libraries such as Vosk, Whisper, Rasa, or similar; just document what you choose and why. Clean architecture, readable code, and future extensibility matter more to me than flashy UI. If you’ve shipped anything comparable—desktop dictation tools, Jarvis-style assistants, or speech-enabled productivity apps—let me know, and point me to a live demo or repo.