AI Wine Expert Voice Assistant

Job ID: 40136416

Budget: €12 – €18 EUR

I’m building a conversational “digital sommelier” that lives both on our website and inside our iOS / Android app. The assistant must fluidly switch between spoken dialogue and text chat, greeting visitors, uncovering their taste profile, and then suggesting bottles that match their palate, meal, or special occasion. Follow-up questions such as grape varietal origins, regional characteristics, ideal serving temperatures, or cellar-aging advice should feel just as natural as speaking to a human wine pro, with answers adjusting in real time as the conversation evolves.

Core functions I expect
• Natural-language understanding and generation for both voice and text
• Real-time speech-to-text and text-to-speech with a pleasant, brand-appropriate voice
• A recommendation engine that factors in stated preferences, food pairing, and occasion context, pulling live data from our wine catalog (REST/GraphQL)
• Memory of prior turns so the agent can refine suggestions without the user repeating themselves
• Simple fallback handling when a request is outside scope, plus graceful escalation to a human chat

Integration environment
The agent will sit inside our React web front end and Flutter mobile app. I’m agnostic on the underlying stack—as long as you can justify it—but I expect a production-ready solution that may combine OpenAI or Rasa for NLU/NLG, Dialogflow CX or a comparable orchestration layer, and a robust TTS/STT service such as Amazon Polly, Google Cloud, or ElevenLabs. Please outline any third-party licensing needs up front.

Deliverables
1. Conversation flow and intent/entity schema
2. Fully coded agent with API hooks to our wine database (Swagger docs available)
3. Web and mobile SDK integration components with sample pages
4. Admin console or config file to tweak prompts, voice, and fallback responses
5. Test scripts plus recorded demo showcasing both modalities
6. Deployment guide and knowledge-transfer session for my development team

Acceptance criteria
• First-response latency < 2 seconds for text; < 3 seconds for voice
• At least 90 % intent recognition accuracy on our test set
• Recommendation relevance rated 4/5 or higher by a panel of ten users
• Zero uncaught errors during a 50-turn stress conversation

If you have shipped end-to-end voice agents before—especially in the food & beverage space—let’s talk.