LLM Prompt Engineering for Natural Dialogue

Job ID: 40157932

Budget: $8 – $15 USD

My outbound voice-AI platform already ingests calls through Twilio, hands audio off to Vapi.ai and LiveKit for real-time streaming, and routes text into OpenAI’s GPT-4 via a Python micro-services layer. The next milestone is tightening the prompt layer so every turn of the conversation sounds genuinely human, shows empathy, and keeps context even across lengthy calls.

The scope here is entirely on the AI & ML stack, with prompt engineering as the critical path. I will provide:

• Existing system prompts, call transcripts, and analytics
• Direct API access to the OpenAI endpoints and my internal orchestration service
• A staging server for rapid, live testing against sample callers

What I expect back:

• A modular prompt library that can be slotted straight into the current Python codebase
• Clearly commented instructions showing where and how each prompt variant should be invoked
• A small evaluation script or notebook that demonstrates response quality against a set of test dialogues, focusing on warmth, clarity, and low hallucination rate
• A brief read-me outlining best practices for future prompt updates

Acceptance criteria: when the new prompts are enabled on staging, at least 90 % of sampled calls must pass a human-likeness rubric (I’ll share the rubric) without sacrificing factual accuracy or latency.

No work is needed on the voice pipeline or backend architecture; everything is already wired. The entire effort is about crafting, testing, and iterating prompts to achieve natural, context-aware dialogue.