Dev for AI/ML Model Training and Fine-Tuning

Job ID: 39481062

Budget: $2 – $8 USD

I’m looking for a skilled and experienced Machine Learning Engineer to help fine-tune and extend the functionality of Microsoft’s Magma-8B multimodal vision-language model. This is a cutting-edge AI development project where your work will be paired with a set of four unique LLaMA-based reasoning models I’ve created. As a key deliverable, you will help build a modular and flexible wrapper around Magma-8B, capable of integrating seamlessly with different reasoning backbones via a simple "backbone" folder structure.

Your main task will be to modify Magma’s visual encoder so that it can visually process PDF documents (converted to images) containing detailed context like system prompts, knowledge references, conversation history, and tool definitions — all without converting this information into raw text, thus preserving token space and formatting integrity. You will also ensure the model can handle UI screenshots and real-world environment images, maintaining performance across varied visual content.

This project will also require building a CLI-based override and training the visual encoder using public Document QA datasets before shifting to a specialized internal dataset. The final model — renamed HiMind-VR — should be delivered fully trained, tested, and documented.

Key Deliverables:

Modify Magma to treat any LLaMA-based model from ./backbone/ as a pluggable text model.
Enable Magma to visually interpret PDF documents (via OCR and pdf2image).
Ensure seamless input merging from PDFs, UI images, and environmental images.
Build CLI override to control backbone selection and configuration.
Train and validate the updated visual wrapper using public DocQA datasets.
Deliver a fully working Magma variant named HiMind-VR with documentation and test results.

Ideal Skills and Experience:

Experience with multimodal models and vision-language integration
Strong understanding of Magma, LLaMA-based models, and modular ML systems
Familiarity with pdf2image, OCR pipelines, and structured data extraction
Prior work on document QA, tool call processing, and reasoning models
Excellent Python skills with tools like PyTorch, HuggingFace Transformers, Selenium (for UI image generation if needed)

Preferred Datasets for Initial Training and Suggested Backbone Stand-in During Training - Full related details and requirements will be provided.

This is a high-impact AI R&D project, ideal for someone who wants to work on the next frontier of AI model architecture and flexible multimodal inference. If you're confident in building scalable ML systems that can understand both visual and textual information, this is your opportunity to shine.

Please include examples of relevant past work, your approach to fine-tuning multimodal models, and your estimated timeframe to complete the key milestones listed.