Vision-Language Model Development for Manga Summarization

Job ID: 39106479

Budget: $10 – $400 USD

I'm seeking an expert to train a vision-language model capable of identifying emotional expressions and actions character names and a lot more in manhwa/manga images. The AI will generate summaries of these pictures, then my code base uses edge tts to generate voice then combine the voice and the respective image into a video. This task includes:

expected output: https://www.youtube.com/watch?v=XiItNbTkPZw&ab_channel=AniManga
https://youtu.be/TPFn4TO2Xa8?si=EJfsteENlG2o1XGC
https://youtu.be/NBj82cT454s?si=yok2AhLMh4dymq7S
https://youtu.be/oaghzvWWPW0?si=ddZ2mV9CLmtQkBJy

- Integrating the trained model into my primary Python and Node.js codebase.
- Debugging and optimizing the existing code.
- Implementing additional features, including a multilingual summary generator.
- Enhancing the output quality using improved image slicing techniques and more.

The ideal candidate will have extensive experience in AI development, particularly with vision-language models, and a strong proficiency in Python and Node.js. Code optimization, bug fixing, and feature addition are integral to this project. Previous work with manga or video content creation will be a significant advantage.