ML enginnering / LLMOps Task - on cloud (hugging-face) -- 2

Job ID: 37756109

Budget: $30 – $250 USD

please note the purpose of the task is ml-engineering focus AND NOT accuracy/ds focus

for example spped up train effiencecy / speed inference cluster train using distributed training/gpu for example/k8s

run the following evaluations using hugging face API.

3 Models: Falcon 7B, Llama 2 7B, Phi-2

6 Tasks: HellaSwag, MMLU, wikitext, Lambada, ARC Challenge, TruthfulQA

Compute Cluster: 4x Nvidia Tesla V100
Your goal is to reach the fastest speed and minimal cost for running all these evaluations on a given compute cluster. You can employ any method that you can think of to accelerate this process (using only a single instance of the provided computing).
The final job you create should include producing all evaluations on all tasks and all models provided. Please provide the following results for your final job:

The total cost of the final run.

The total time it took the run to complete.

A written report on the developed infrastructure. (azure env prefered)
Related categories: Machine Learning (ML) NLP Hugging Face MLOps