Set Up llama.cpp with GPU (RTX 5080) and Optimize for Mixtral – Test Q5_K_M vs Q4_K_M

Job ID: 39332220

Budget: $30 – $250 AUD

I need a skilled developer to remotely configure my Windows PC for optimal use with the Mixtral model via llama.cpp.

The goal is to install everything needed, build llama.cpp with full CUDA GPU acceleration for my RTX 5080, and optimize the setup so it runs future Mixtral .gguf models with maximum speed and stability.

Key Tasks:
Install all required dependencies (build tools, Python, latest CUDA)

Build llama.cpp with LLAMA_CUBLAS=1 for GPU support

Optimize runtime settings for my RTX 5080 (context size, threads, GPU layers)

Test two quant formats: Q4_K_M and Q5_K_M

Compare performance of both (speed, GPU load, stability)

Show nvidia-smi during test to confirm GPU is working

Leave the system ready for me to drop in Mixtral later and run instantly

Requirements:
You must use AnyDesk or RustDesk only — no long-term access

You’ll be supervised during the session

Please explain what you’re doing as you go

Clean, efficient install — no extra bloat or services

Deliverables:
llama.cpp built and ready with GPU

My system fully prepared for future Mixtral usage

Performance test results and confirmation of GPU use

Looking for someone with experience in CUDA, GPU optimization, and LLM model deployment. Please be ready to recommend ideal Mixtral quant + config.