Multimodal LLM Fine-Tuning Expert
Budget: $50 – $0 USD
I have an in-house data that blends text, video and image content and I am ready to push a GPT-class transformer beyond pure language. What I need is a specialist who already speaks the language of DPO and GRPO and can translate those techniques into a practical training pipeline. The objective is straightforward: take an existing open-weight model, fine-tune it on my multimodal set, and return a checkpoint that outperforms vanilla GPT on the tasks my team actually cares about.
You will be free to choose between TensorFlow or PyTorch for the heavy lifting—both are wired into our environment—so long as the final codebase is reproducible on standard CUDA hardware. The data are pre-sharded; you will focus on building the loaders, aligning the modalities, and steering the optimization with preference-based feedback. I will provide the raw assets, any schema documentation, and access to a staging GPU cluster.
Deliverables
• End-to-end training scripts (command-line runnable)
• Trained weights and config files compatible with Hugging Face transformers
• A concise README that explains dataset preprocessing, hyper-parameters and how to resume or extend training
Acceptance criteria
• Model ingests text, video frames and images in a single forward pass
• Demonstrated improvement over baseline on an internal test set we will share (BLEU for text, retrieval accuracy for visuals)
• Training run can be reproduced from scratch using the provided scripts on a single A100 in < 48h
If you have recent hands-on results with preference optimization for LLMs and you are comfortable juggling mixed-modality tensors, let’s talk.
You will be free to choose between TensorFlow or PyTorch for the heavy lifting—both are wired into our environment—so long as the final codebase is reproducible on standard CUDA hardware. The data are pre-sharded; you will focus on building the loaders, aligning the modalities, and steering the optimization with preference-based feedback. I will provide the raw assets, any schema documentation, and access to a staging GPU cluster.
Deliverables
• End-to-end training scripts (command-line runnable)
• Trained weights and config files compatible with Hugging Face transformers
• A concise README that explains dataset preprocessing, hyper-parameters and how to resume or extend training
Acceptance criteria
• Model ingests text, video frames and images in a single forward pass
• Demonstrated improvement over baseline on an internal test set we will share (BLEU for text, retrieval accuracy for visuals)
• Training run can be reproduced from scratch using the provided scripts on a single A100 in < 48h
If you have recent hands-on results with preference optimization for LLMs and you are comfortable juggling mixed-modality tensors, let’s talk.