Python LLM Performance Testing Expert -- 6
Budget: $30 – $250 USD
I am in need of a Python expert who can perform a deep dive into the performance of various LLMs under different conditions. The main focus of this task is to benchmark different LLMs and evaluate their performance against each other.
Key responsibilities:
- Conducting performance tests on LLMs.
- Analysing the results and comparing them to identify the best performing system.
- Providing recommendations on how to optimize the performance of the chosen LLM.
- Set up a local testing environment (chatbot, playground dashboard)
- Set up a cloud server (API server)
- Google Colab, Jupyter Notebook
- All configuration, test scripts in python, bash and exe)
- Dockerized images for future implementation
- Our Korean dataset of 20,000 published to HuggingFace workspace based on an existing Korean dataset
- documentation 1. testing items/procedures 2. instructions, manuals of this project.
You will also be required to integrate the API server with third-party APIs. Experience in this area is highly beneficial.
Ideal Skills:
- Strong expertise in Python programming.
- Previous experience in performance testing, specifically LLMs.
- Familiarity with benchmarking tools and methodologies.
- Ability to work with APIs and integrate them effectively.
Models to consider:
llama3-8b-8192
text-davinci-003
Mixtral 8x22b
Whisper
Our datasets on hf, /datasets/uconcreative/slmDatasets
Reference:
https://pf7.eggs.or.kr/aigenerative_overview.html
Sample testing items (send me a permission request with your name plz)
https://docs.google.com/document/d/16GFgABbbnH7RXgL56SXNu41xRspstgtVcy0TQxmwd-A/edit
Project term: 4 days
Key responsibilities:
- Conducting performance tests on LLMs.
- Analysing the results and comparing them to identify the best performing system.
- Providing recommendations on how to optimize the performance of the chosen LLM.
- Set up a local testing environment (chatbot, playground dashboard)
- Set up a cloud server (API server)
- Google Colab, Jupyter Notebook
- All configuration, test scripts in python, bash and exe)
- Dockerized images for future implementation
- Our Korean dataset of 20,000 published to HuggingFace workspace based on an existing Korean dataset
- documentation 1. testing items/procedures 2. instructions, manuals of this project.
You will also be required to integrate the API server with third-party APIs. Experience in this area is highly beneficial.
Ideal Skills:
- Strong expertise in Python programming.
- Previous experience in performance testing, specifically LLMs.
- Familiarity with benchmarking tools and methodologies.
- Ability to work with APIs and integrate them effectively.
Models to consider:
llama3-8b-8192
text-davinci-003
Mixtral 8x22b
Whisper
Our datasets on hf, /datasets/uconcreative/slmDatasets
Reference:
https://pf7.eggs.or.kr/aigenerative_overview.html
Sample testing items (send me a permission request with your name plz)
https://docs.google.com/document/d/16GFgABbbnH7RXgL56SXNu41xRspstgtVcy0TQxmwd-A/edit
Project term: 4 days