AI Tools Evaluation Pipeline

Job ID: 39531024

Budget: $750 – $1,500 USD

I'm looking for a freelancer to build an evaluation pipeline for Copilot-style tools (e.g., ChatGPT, Microsoft Copilot, ..,). The pipeline should measure:
- Accuracy and groundedness
- Latency and throughput
- Output quality (fluency, helpfulness, conciseness)

Ideal experience:
- ML model evaluation
- Building QA pipelines for AI tools
- Familiarity with benchmarks like BEGIN, GaRAGe, TruthfulQA

Please include relevant experience in your bid.