LLM API Cost Optimization
Budget: $750 – $1,500 USD
I run a Telegram bot that intelligently switches between GPT-4, Claude 3, and a few additional LLMs. It delivers great answers, but each call is becoming expensive. I need an engineer (or small team) who can drive those costs down without letting response quality slip—high-quality replies are non-negotiable for my users.
Here’s what I’m looking for:
• Conduct a detailed audit of current GPT, Claude and other-LLM usage inside the bot (token counts, temperature settings, system vs. user prompt length, retry logic, etc.).
• Propose and implement concrete savings tactics—prompt trimming, token budgeting, dynamic model selection, response caching, batching, or any creative approach you’ve proven elsewhere.
• Instrument quality checks so we can A/B test every tweak and confirm answers stay at the same high standard.
• Leave me with clean, well-commented code and a short memo outlining what was changed, why, and how to maintain it.
Acceptance criteria
– Minimum 70 % reduction in average cost per successful request measured over a representative 7-day window.
– No statistically significant drop in user-rated answer quality compared with the current baseline.
Please highlight prior experience doing cost or performance optimisation on GPT, Claude, or comparable LLM APIs when you reply.
Here’s what I’m looking for:
• Conduct a detailed audit of current GPT, Claude and other-LLM usage inside the bot (token counts, temperature settings, system vs. user prompt length, retry logic, etc.).
• Propose and implement concrete savings tactics—prompt trimming, token budgeting, dynamic model selection, response caching, batching, or any creative approach you’ve proven elsewhere.
• Instrument quality checks so we can A/B test every tweak and confirm answers stay at the same high standard.
• Leave me with clean, well-commented code and a short memo outlining what was changed, why, and how to maintain it.
Acceptance criteria
– Minimum 70 % reduction in average cost per successful request measured over a representative 7-day window.
– No statistically significant drop in user-rated answer quality compared with the current baseline.
Please highlight prior experience doing cost or performance optimisation on GPT, Claude, or comparable LLM APIs when you reply.