AI Lyrics Timing Windows Software
Budget: $250 – $750 USD
I need a Windows-based application that can take an audio track and, with the help of AI, place every lyric in perfect sequence on a timeline. The core requirements are:
• Automatic timing: the software should listen to the song, recognise each word or syllable, and stamp accurate in/out points.
• Gender detection: when the voice is male the lyric must appear in one colour, and when the voice is female it must appear in a different colour. Using distinct text colours (not bold or background changes) is the chosen method.
• Real-time and manual editing: while the AI does the heavy lifting, I still want to scrub through the track and tweak individual timecodes or lyric blocks instantly if anything is off.
• Language handling: the engine has to work flawlessly in English first, with the architecture left open for additional languages later on.
A simple, intuitive timeline view—similar to most DAWs or karaoke editors—will make the manual adjustment workflow painless. If you are comfortable combining speech-to-text, diarisation (for gender), and a responsive UI in a Windows desktop environment, I’d love to see how you would approach this.
Final deliverable: an installer or portable build, plus source code and a brief README so I can compile or extend the project in future.
• Automatic timing: the software should listen to the song, recognise each word or syllable, and stamp accurate in/out points.
• Gender detection: when the voice is male the lyric must appear in one colour, and when the voice is female it must appear in a different colour. Using distinct text colours (not bold or background changes) is the chosen method.
• Real-time and manual editing: while the AI does the heavy lifting, I still want to scrub through the track and tweak individual timecodes or lyric blocks instantly if anything is off.
• Language handling: the engine has to work flawlessly in English first, with the architecture left open for additional languages later on.
A simple, intuitive timeline view—similar to most DAWs or karaoke editors—will make the manual adjustment workflow painless. If you are comfortable combining speech-to-text, diarisation (for gender), and a responsive UI in a Windows desktop environment, I’d love to see how you would approach this.
Final deliverable: an installer or portable build, plus source code and a brief README so I can compile or extend the project in future.