Script for Multi-Source Prompt Benchmark

Job ID: 38277464

Budget: $30 – $60 USD

I'm seeking a proficient python and genAI specialist to wwrite a script to benchmark prompt responses from manuall labelling, response from custom prompt, GPT4, GPT3.5, Mistral  using F1-score, precision and recall.

The input will be :
1. pdf document
2. set of questions
3. set of manually extracted answers for the questions
4. set of responses from custom prompt

logic to:
create prompt for GPT4, GPT3.5, Mistral to ask question that are provided in the input from the document
Calculate F1-score, precision and recall.

output
Write questions, answers for manually extracted responses, response from GPT4, GPT3.5, Mistralfor F1-score, precision and recall in an excel or csv file
Related categories: Python GenAI