Enhance AI Video App & Integrate Musa TTS
Budget: $10 – $30 USD
Project Title
Developer Needed to Improve AI History Video Generation App + Add Musa TTS Audio API Integration
Project Description
I already have an AI-powered history video generation app built using Google AI Studio / AI workflow. The app generates history-style videos using scripts, visual descriptions, images/assets, and narration.
I am looking for an experienced developer who can fix current issues, improve the video generation workflow, integrate my custom audio generation API, and align the final video output with a reference YouTube video style.
Reference video style:
https://www.youtube.com/watch?v=n1ZIcmXI5SU&t=21s
The goal is not to rebuild everything from scratch unless absolutely required. I want the existing app to be improved, optimized, and upgraded so it can produce better historical documentary-style videos.
Main Problems to Fix
• The images/assets fetched by the app are not accurately matching the visual descriptions.
• Some visuals are not historically accurate or relevant to the scene.
• The app needs better prompt logic for converting script scenes into accurate visual search/image generation descriptions.
• The final video style needs to feel closer to the reference.
• Some fetched videos are not loaded successfully
Required Work
• Review/audit the current AI video generation app and identify what needs to be fixed.
• Improve the image/asset fetching system so visuals match the scene description more accurately.
• Improve historical accuracy of the selected images and visual assets.
• Add more free sources/APIs for historical stock assets, public domain images, museum/archive images, and historical visuals.
• Improve the prompt workflow for scene-by-scene visual generation.
• Help align the video output with the reference YouTube style, including:
o documentary-style pacing
o historical mood and atmosphere
o better scene-to-image matching
o smooth flow from script to visuals
o suitable narration timing
o more professional history-video structure
• Integrate my own audio generation tool called Musa TTS.
• Musa TTS already has:
o cloned voice
o local server
o API key
• Set up ngrok so the local Musa TTS server can be accessed by the Google AI Studio / AI video generation app.
• Connect the app to the Musa TTS API using API key authentication.
• Make sure narration/audio generation works automatically inside the video generation pipeline.
• Test the full workflow:
o script generation
o scene splitting
o visual description generation
o accurate historical image/asset fetching
o Musa TTS voice generation
o final video assembly/export
• Fix bugs related to APIs, image fetching, prompts, audio generation, or app workflow.
Reference Style Requirement
I will provide a reference YouTube video/channel style. I want the app output to be inspired by that style and improved in that direction.
The developer should analyze the reference and help make the app generate videos with a similar video editing style.
The final output should feel polished, cinematic, and suitable for history storytelling content.
Developer Requirements
• Experience with AI apps and API integrations.
• Experience with Google AI Studio / Gemini API or similar AI tools.
• Experience with backend API connections.
• Experience with local server to cloud connection using ngrok.
• Experience with API key authentication and secure endpoint setup.
• Experience with image search APIs, stock asset APIs, or public domain image sources.
• Ability to debug existing apps instead of only building from scratch.
• Understanding of AI video generation workflows.
• Bonus if you have experience with:
o TTS systems
o voice cloning tools
o AI documentary/history video tools
o public domain archive/museum image APIs
o video automation pipelines
Final Goal
The final goal is to make my AI history video generation app produce more accurate, professional, documentary-style historical videos with better visuals and high-quality cloned voice narration using Musa TTS.
Developer Needed to Improve AI History Video Generation App + Add Musa TTS Audio API Integration
Project Description
I already have an AI-powered history video generation app built using Google AI Studio / AI workflow. The app generates history-style videos using scripts, visual descriptions, images/assets, and narration.
I am looking for an experienced developer who can fix current issues, improve the video generation workflow, integrate my custom audio generation API, and align the final video output with a reference YouTube video style.
Reference video style:
https://www.youtube.com/watch?v=n1ZIcmXI5SU&t=21s
The goal is not to rebuild everything from scratch unless absolutely required. I want the existing app to be improved, optimized, and upgraded so it can produce better historical documentary-style videos.
Main Problems to Fix
• The images/assets fetched by the app are not accurately matching the visual descriptions.
• Some visuals are not historically accurate or relevant to the scene.
• The app needs better prompt logic for converting script scenes into accurate visual search/image generation descriptions.
• The final video style needs to feel closer to the reference.
• Some fetched videos are not loaded successfully
Required Work
• Review/audit the current AI video generation app and identify what needs to be fixed.
• Improve the image/asset fetching system so visuals match the scene description more accurately.
• Improve historical accuracy of the selected images and visual assets.
• Add more free sources/APIs for historical stock assets, public domain images, museum/archive images, and historical visuals.
• Improve the prompt workflow for scene-by-scene visual generation.
• Help align the video output with the reference YouTube style, including:
o documentary-style pacing
o historical mood and atmosphere
o better scene-to-image matching
o smooth flow from script to visuals
o suitable narration timing
o more professional history-video structure
• Integrate my own audio generation tool called Musa TTS.
• Musa TTS already has:
o cloned voice
o local server
o API key
• Set up ngrok so the local Musa TTS server can be accessed by the Google AI Studio / AI video generation app.
• Connect the app to the Musa TTS API using API key authentication.
• Make sure narration/audio generation works automatically inside the video generation pipeline.
• Test the full workflow:
o script generation
o scene splitting
o visual description generation
o accurate historical image/asset fetching
o Musa TTS voice generation
o final video assembly/export
• Fix bugs related to APIs, image fetching, prompts, audio generation, or app workflow.
Reference Style Requirement
I will provide a reference YouTube video/channel style. I want the app output to be inspired by that style and improved in that direction.
The developer should analyze the reference and help make the app generate videos with a similar video editing style.
The final output should feel polished, cinematic, and suitable for history storytelling content.
Developer Requirements
• Experience with AI apps and API integrations.
• Experience with Google AI Studio / Gemini API or similar AI tools.
• Experience with backend API connections.
• Experience with local server to cloud connection using ngrok.
• Experience with API key authentication and secure endpoint setup.
• Experience with image search APIs, stock asset APIs, or public domain image sources.
• Ability to debug existing apps instead of only building from scratch.
• Understanding of AI video generation workflows.
• Bonus if you have experience with:
o TTS systems
o voice cloning tools
o AI documentary/history video tools
o public domain archive/museum image APIs
o video automation pipelines
Final Goal
The final goal is to make my AI history video generation app produce more accurate, professional, documentary-style historical videos with better visuals and high-quality cloned voice narration using Musa TTS.
Related categories:
Full Stack Development
API Integration
AI Model Integration
AI Development
AI Video
Lovable