Adapt Bittensor CLM Model Tuning script to work with multiple GPUs
Budget: €30 – €250 EUR
I am looking for someone to convert the Bittensor CLM finetuning script to support multi gpu. The bittensor version of the script has been adapted from Hugging Face's transformers/language-modeling code and can be found here: https://github.com/opentensor/clm_model_tuning. The script uses the datasets hosted on the IPFS Genesis dataset from Bittensor.
Currently the script does not work with GPU parallelism and I need someone to make it work with 2/4/6/8 GPUs in parallel to be able to train larger models or train them faster. Models which I want to finetune are https://huggingface.co/EleutherAI/gpt-neo-1.3B, https://huggingface.co/EleutherAI/gpt-neo-2.7B and https://huggingface.co/EleutherAI/gpt-j-6B. The finetuning script should reduce the loss, support multi gpu, train on the bittensor dataset and work for the huggingface models mentioned above.
More information about the Bittensor CLM script: https://docs.bittensor.com/nested/FineTuning.html
More information about the dataset: https://docs.bittensor.com/nested/TheDataset.html
Currently the script does not work with GPU parallelism and I need someone to make it work with 2/4/6/8 GPUs in parallel to be able to train larger models or train them faster. Models which I want to finetune are https://huggingface.co/EleutherAI/gpt-neo-1.3B, https://huggingface.co/EleutherAI/gpt-neo-2.7B and https://huggingface.co/EleutherAI/gpt-j-6B. The finetuning script should reduce the loss, support multi gpu, train on the bittensor dataset and work for the huggingface models mentioned above.
More information about the Bittensor CLM script: https://docs.bittensor.com/nested/FineTuning.html
More information about the dataset: https://docs.bittensor.com/nested/TheDataset.html