Real-Time Speaking Feedback Tool
Budget: €30 – €250 EUR
I’m a YouTuber who wants immediate, practical coaching while I’m recording or live-streaming. The goal is a desktop application that lets me pick any connected microphone and camera, sends the live feed to Gemini (or the most effective alternative you recommend), and listens for moments where I can improve myself for the rhythms or anything else, listener-friendly rhythm. When that happens, the app should break could speak to me so I can correct myself on the spot.
I work on both macOS and Windows.
So I’ll let you decide what is the best.
I’m happy with an Electron wrapper, a native Swift + .NET pair, or any framework you believe will keep latency low; just outline why you think it’s the best route. If the Gemini Live API is not yet public or stable, feel free to integrate an equivalent model, but please keep the code modular so I can swap providers later.
Acceptance criteria:
• I can choose the audio and video sources from a simple dropdown.
• The app starts analysing instantly, with no perceptible lag in the monitor feed.
• Whenever I speed up or slow down beyond a definable threshold, I hear a clear audio cue.
• Settings let me adjust sensitivity and the style.
• You supply complete, well-commented source code and a quick setup guide
If you’ve built low-latency AI tools, especially with Gemini, Whisper, or WebRTC, let me know; I’m keen to move quickly once we agree on an approach.
I work on both macOS and Windows.
So I’ll let you decide what is the best.
I’m happy with an Electron wrapper, a native Swift + .NET pair, or any framework you believe will keep latency low; just outline why you think it’s the best route. If the Gemini Live API is not yet public or stable, feel free to integrate an equivalent model, but please keep the code modular so I can swap providers later.
Acceptance criteria:
• I can choose the audio and video sources from a simple dropdown.
• The app starts analysing instantly, with no perceptible lag in the monitor feed.
• Whenever I speed up or slow down beyond a definable threshold, I hear a clear audio cue.
• Settings let me adjust sensitivity and the style.
• You supply complete, well-commented source code and a quick setup guide
If you’ve built low-latency AI tools, especially with Gemini, Whisper, or WebRTC, let me know; I’m keen to move quickly once we agree on an approach.