Vision-Language Model Training for Manga Summarization

Job ID: 39103506

Budget: $30 – $250 USD

I'm seeking an expert in AI, Node.js, and Python to train a vision-language model tailored to my needs. The purpose of this model is to analyze images from manhwa/manga and generate detailed summaries post-training. Once the model meets my specifications, you'll be responsible for integrating it into my main codebase. and add a multilingual feature to generate summaries in many different languages debug it so I can run it without any error and make some changes to the main codebase.

expected output: https://www.youtube.com/watch?v=XiItNbTkPZw&ab_channel=AniManga
https://youtu.be/TPFn4TO2Xa8?si=EJfsteENlG2o1XGC

the summary generated by the ai has to match these video summary

Key Responsibilities:
- Train a vision-language model for detailed manga/manhwa summarization.
- Analyze mixed content within the series, including characters, dialogues, and action scenes.
- Process input images in various formats, primarily webp, jpg, and png.
- Install the finalized model into my primary codebase.

Ideal Skills and Experience:
- Proficiency in AI, particularly in training vision-language models.
- Strong skills in Node.js and Python.
- Experience working with mixed visual content.
- Familiarity with image formats including webp, jpg, and png.
- Prior experience in integrating AI models into codebases.