LLM API Cost Optimization

Job ID: 39784955

Budget: $750 – $1,500 USD

I run a Telegram bot that intelligently switches between GPT-4, Claude 3, and a few additional LLMs. It delivers great answers, but each call is becoming expensive. I need an engineer (or small team) who can drive those costs down without letting response quality slip—high-quality replies are non-negotiable for my users.

Here’s what I’m looking for:

• Conduct a detailed audit of current GPT, Claude and other-LLM usage inside the bot (token counts, temperature settings, system vs. user prompt length, retry logic, etc.).
• Propose and implement concrete savings tactics—prompt trimming, token budgeting, dynamic model selection, response caching, batching, or any creative approach you’ve proven elsewhere.
• Instrument quality checks so we can A/B test every tweak and confirm answers stay at the same high standard.
• Leave me with clean, well-commented code and a short memo outlining what was changed, why, and how to maintain it.

Acceptance criteria
– Minimum 70 % reduction in average cost per successful request measured over a representative 7-day window.
– No statistically significant drop in user-rated answer quality compared with the current baseline.

Please highlight prior experience doing cost or performance optimisation on GPT, Claude, or comparable LLM APIs when you reply.