AI Evaluation & Red Teaming Tool Development

Job ID: 40447985

Budget: $30 – $250 CAD

Requirements for the AI Evaluation & Voice Testing Platform,
Phase 1 — Voice Load Testing & Core Evaluation Platform

This phase will include:

• Test Suite Dashboard (create/manage evaluation suites)
• SIP / API / Webhook connection modes
• Voice load testing framework (SIPp based)
• Concurrent call simulation (demo scale locally, scalable to 3000 ports on server)
• Deterministic flows (scripted IVR tests)
• Agentic flows using LLM for dynamic conversations
• Retry logic for failed calls
• Technical metrics collection (latency, success rate, call failures)
• Basic reporting dashboard

Phase 2 — AI Evaluation Engine & Red Teaming

Includes:

• AI evaluation scoring (intent accuracy, entity extraction)
• Hallucination detection
• Prompt-injection / red-team testing scenarios
• AI judge using LLM (OpenAI / Vertex AI / Bedrock)
• Conversation transcript analysis
• Evaluation scorecards per test run

Advanced Observability & Reporting

Includes:

• Grafana integration and open telemetry
• Full analytics dashboards
• Test result comparison between versions
• Performance regression detection
• Exportable reports for stakeholders



We can start with Phase 1 (Voice Load Testing + Core Evaluation) and expand the platform gradually.