Self-Hosted LLM Prompt Migration

Job ID: 40486338

Budget: £10 – £15 GBP

I’m moving our AI inference off OpenAI and onto a Tesla P100 16 GB box that already runs qwen2.5 7B/14B on Ollama. The backend is wired to switch between local and remote per-prompt, so the infrastructure work is minimal; the real task is model selection, tuning and validation.

Cost reduction is the driving motive, but I will only flip a prompt when the dashboards show zero drop in the accuracy of CV match scores (real interview rate is the secondary check). We have four production prompts; each will run in shadow mode until its metrics are indistinguishable from the current OpenAI baseline.

What I need from you
• Pick or fine-tune the best qwen2.5 variant—or another model that will fit and perform on a single P100—then set quantisation, context window and batching so latency stays reasonable.
• Run shadow tests, analyse the match-score deltas and recommend when we should switch traffic.
• Repeat for all four prompts, one at a time, over roughly 6–10 weeks at 15–25 hours per week.

To apply, tell me:
1. Your hourly rate.
2. One similar migration or on-prem LLM project you’ve shipped.
3. The first change you’d make given a P100 and qwen2.5.

If you thrive on squeezing the most out of limited GPU memory while keeping quality rock solid, I’d love to hear your plan.