Optimize DeepSeek 7B Model Training
Budget: ₹1,500 – ₹12,500 INR
I'm seeking an expert in LLM and DeepSpeed to help fine-tune our DeepSeek 7B base model for a multi-label classification task. We are using HuggingFace transformers version 4.40.0, DeepSpeed 0.14.0, and a setup with 4x H200 GPUs, each with 140GB VRAM. The project is underway, but we need immediate assistance to address several critical issues.
Key Requirements:
- Resolve the issue of models not loading onto GPUs, as current GPU usage is stuck at 0%.
- Optimize the use of TrainingArguments, including evaluation strategy and other parameters.
- Review and enhance our DeepSpeed configuration for performance and memory optimization.
- Ensure full utilization and parallelism across all 4 GPUs to enable efficient multi-GPU training.
- Assist in achieving scalable training without crashes or idling.
The dataset is preprocessed, tokenized, and ready for use, consisting of approximately 1.3 million samples in JSONL format. The codebase is mostly functional but requires debugging and tuning for optimal performance.
Resources Provided:
- Access to a remote instance with 4x H200 GPUs (561 GB total VRAM), 192 vCPUs, and 1 TB RAM.
- Complete access to the codebase and dataset.
Ideal Skills and Experience:
- Expertise in DeepSpeed and HuggingFace transformers.
- Strong understanding of GPU optimization and parallelism.
- Experience with AdamW optimizer and debugging model loading issues.
- Proficiency in performance tuning and memory management for large-scale model training.
Your expertise will be crucial in ensuring that our model training process is efficient and fully utilizes the available resources. If you have the skills and experience to tackle these challenges, I look forward to your bid.
Key Requirements:
- Resolve the issue of models not loading onto GPUs, as current GPU usage is stuck at 0%.
- Optimize the use of TrainingArguments, including evaluation strategy and other parameters.
- Review and enhance our DeepSpeed configuration for performance and memory optimization.
- Ensure full utilization and parallelism across all 4 GPUs to enable efficient multi-GPU training.
- Assist in achieving scalable training without crashes or idling.
The dataset is preprocessed, tokenized, and ready for use, consisting of approximately 1.3 million samples in JSONL format. The codebase is mostly functional but requires debugging and tuning for optimal performance.
Resources Provided:
- Access to a remote instance with 4x H200 GPUs (561 GB total VRAM), 192 vCPUs, and 1 TB RAM.
- Complete access to the codebase and dataset.
Ideal Skills and Experience:
- Expertise in DeepSpeed and HuggingFace transformers.
- Strong understanding of GPU optimization and parallelism.
- Experience with AdamW optimizer and debugging model loading issues.
- Proficiency in performance tuning and memory management for large-scale model training.
Your expertise will be crucial in ensuring that our model training process is efficient and fully utilizes the available resources. If you have the skills and experience to tackle these challenges, I look forward to your bid.