Self-Hosted LLM Prompt Migration
Budget: £10 – £15 GBP
I’m moving our AI inference off OpenAI and onto a Tesla P100 16 GB box that already runs qwen2.5 7B/14B on Ollama. The backend is wired to switch between local and remote per-prompt, so the infrastructure work is minimal; the real task is model selection, tuning and validation.
Cost reduction is the driving motive, but I will only flip a prompt when the dashboards show zero drop in the accuracy of CV match scores (real interview rate is the secondary check). We have four production prompts; each will run in shadow mode until its metrics are indistinguishable from the current OpenAI baseline.
What I need from you
• Pick or fine-tune the best qwen2.5 variant—or another model that will fit and perform on a single P100—then set quantisation, context window and batching so latency stays reasonable.
• Run shadow tests, analyse the match-score deltas and recommend when we should switch traffic.
• Repeat for all four prompts, one at a time, over roughly 6–10 weeks at 15–25 hours per week.
To apply, tell me:
1. Your hourly rate.
2. One similar migration or on-prem LLM project you’ve shipped.
3. The first change you’d make given a P100 and qwen2.5.
If you thrive on squeezing the most out of limited GPU memory while keeping quality rock solid, I’d love to hear your plan.
Cost reduction is the driving motive, but I will only flip a prompt when the dashboards show zero drop in the accuracy of CV match scores (real interview rate is the secondary check). We have four production prompts; each will run in shadow mode until its metrics are indistinguishable from the current OpenAI baseline.
What I need from you
• Pick or fine-tune the best qwen2.5 variant—or another model that will fit and perform on a single P100—then set quantisation, context window and batching so latency stays reasonable.
• Run shadow tests, analyse the match-score deltas and recommend when we should switch traffic.
• Repeat for all four prompts, one at a time, over roughly 6–10 weeks at 15–25 hours per week.
To apply, tell me:
1. Your hourly rate.
2. One similar migration or on-prem LLM project you’ve shipped.
3. The first change you’d make given a P100 and qwen2.5.
If you thrive on squeezing the most out of limited GPU memory while keeping quality rock solid, I’d love to hear your plan.