Fine tune Llava model on VQA pairs using NIH XRAY images -- 2

Job ID: 39672401

Budget: ₹600 – ₹1,500 INR

The Goal very simple: to fine-tune a vision-language model (LLaVA or CogVLM2) on a meaningful(but as small as feasibnle) subset of the NIH ChestX-ray14 dataset. The goal is to build a system capable of Visual Question Answering (VQA) tailored to medical diagnostics.

Dataset: NIH ChestX-ray14

Task: Fine-tune the model using structured VQA pairs (image + question + answer). We can annotate the data using an LLM or I can provide you with the annotated data since I have already done that.

Model: LLaVA or CogVLM2

Output: A fine-tuned model capable of zero/few-shot inference on unseen X-rays.

Look, I don't need this to be massive. This is just a small submission task needed for my college and the only reason i am posting this here is that I am on a macbook so VLMs are just not possible to be fine tuned locally for me.