AI-Based Script-to-Song Video System Creation

Job ID: 38998837

Budget: $30 – $250 USD

AI Developer Needed to Create a Script-to-Song System Using TV Show Scripts and Video Clips

Please take a look at this example video to get an idea of the style we’re aiming for:

https://www.youtube.com/watch?v=Vovmo1KSW8Y

Description:

We are seeking an experienced AI developer to build a unique system that can take full-length episodes of The Sopranos (or any TV show with available transcripts/subtitles) and create a musical song by extracting iconic lines, making them rhyme, and stitching together video clips from the show. The system should automatically perform the following steps:

Key Requirements:
Transcript and Subtitle Extraction:

The system needs to process video episodes of The Sopranos and extract corresponding dialogue from subtitle files or generate it through speech-to-text transcription (e.g., using OpenAI Whisper or Google Cloud STT).

Line Selection and Rhyme Generation:

The system must analyze the extracted dialogue and automatically select iconic or meaningful lines from the show based on keywords or themes.
It should rewrite or rearrange the selected lines into rhyming lyrics using an AI-powered rhyme generator (e.g., Datamuse API or GPT-4).

Video Clip Matching:

The system must match the selected dialogue lines to specific timestamps in the video episodes and extract those video segments.
Use video processing tools (e.g., FFmpeg) to automatically extract clips based on the identified timestamps.

Song Structure Creation:

The AI should arrange the rhyming lines into a song structure (e.g., verses, chorus, bridge) with proper rhythm and flow.
Use prosody analysis to ensure the lines are rhythmic and fit a musical pattern.

Video Splicing and Editing:

The system must splice the video clips together in the correct sequence based on the song structure.
It should also ensure smooth transitions between clips and adjust lip-syncing (if necessary) using tools like Wav2Lip to make the video appear natural.

Music Composition:

The system must generate a background music track for the song that aligns with the tone and mood of the lyrics.
You can use AI tools like AIVA, MuseNet, or similar music generation platforms.

Final Video Rendering:

The system should render the final video, combining the video clips, song, and music into a cohesive music video.
The video should be polished with optional quality enhancement using tools like Topaz Video AI.

Skills and Experience Required:

Expertise in NLP (Natural Language Processing) for text analysis and rhyme generation.
Familiarity with speech-to-text (STT) transcription and timestamp mapping.
Experience with video processing and tools like FFmpeg.
Knowledge of AI music composition using platforms like AIVA or MuseNet.
Experience with lip-sync correction tools such as Wav2Lip.
Ability to create automated workflows and custom scripts for video editing and synchronization.

Project Scope:

Budget: Please provide your estimate based on the above requirements.
Timeline: Ideally, we would like to see a working prototype in [X weeks/months].

Additional Notes:

We can provide video episodes and subtitle files of The Sopranos for testing.

You should have access to the necessary tools for video editing and music composition, or be willing to work with available APIs and libraries.

If you have experience in AI, NLP, video editing automation, and music generation, we’d love to hear from you!