Modular AI-Powered Media Automation System (GUI + Python + AI Integration)
Budget: $250 – $750 USD
I'm seeking a skilled full-stack developer (or team) to create a modular AI-assisted media automation system, capable of processing content such as Manga, Anime, and Manhwa. The system should perform various tasks including, but not limited to:
https://docs.google.com/document/d/1F1x-V70aOL6gl8YCgmZkjm9N8N3dpYrIQF8P0h9oGTo/edit?usp=sharing
- Coloring
- Summarization
- OCR (Optical Character Recognition)
- Translation
- TTS (Text-to-Speech)
- Slicing
- Video Generation
Key Features:
System Architecture:
Manual step-triggering (no auto-chaining unless run explicitly)
Each module is independent — must run correctly when valid input is present
Steps can be skipped, rearranged, or repeated without error
Full project persistence — resumes from where it left off
Dynamic input/output folder structure
Workflow Modules (Each Can Be Triggered Independently):
Step 1: Input Type Selector: Manga / Manhwa / Anime / Other
Step 2: Media Input (Image or Video)
Image Processing Modules:
Merging
Slicing (e.g. panel splitter)
OCR (Magi V2 or Mistral AI depending on content)
Cleaning (e.g. BubbleBlaster for text bubble removal)
Coloring (Manga only)
Text AI Modules:
Vision model summary (Gemini/GPT)
Text refining (shorten, improve)
Translation (DeepSeek/local model)
Text-to-speech (natural, not robotic)
Video Workflow:
Anime Script Generator (Generates clip script & timestamps from input video)
Anime Cutter (Cuts clip from original anime using timestamps)
Final video creation (YouTube-long style or TikTok-short)
Optional panel animation (zoom, pan, float effects toggleable)
Thumbnail generator (theme-based + prompt to image model)
Clip shortener for TikTok-style output
https://docs.google.com/document/d/1F1x-V70aOL6gl8YCgmZkjm9N8N3dpYrIQF8P0h9oGTo/edit?usp=sharing
- Coloring
- Summarization
- OCR (Optical Character Recognition)
- Translation
- TTS (Text-to-Speech)
- Slicing
- Video Generation
Key Features:
System Architecture:
Manual step-triggering (no auto-chaining unless run explicitly)
Each module is independent — must run correctly when valid input is present
Steps can be skipped, rearranged, or repeated without error
Full project persistence — resumes from where it left off
Dynamic input/output folder structure
Workflow Modules (Each Can Be Triggered Independently):
Step 1: Input Type Selector: Manga / Manhwa / Anime / Other
Step 2: Media Input (Image or Video)
Image Processing Modules:
Merging
Slicing (e.g. panel splitter)
OCR (Magi V2 or Mistral AI depending on content)
Cleaning (e.g. BubbleBlaster for text bubble removal)
Coloring (Manga only)
Text AI Modules:
Vision model summary (Gemini/GPT)
Text refining (shorten, improve)
Translation (DeepSeek/local model)
Text-to-speech (natural, not robotic)
Video Workflow:
Anime Script Generator (Generates clip script & timestamps from input video)
Anime Cutter (Cuts clip from original anime using timestamps)
Final video creation (YouTube-long style or TikTok-short)
Optional panel animation (zoom, pan, float effects toggleable)
Thumbnail generator (theme-based + prompt to image model)
Clip shortener for TikTok-style output