Script for Multi-Source Prompt Benchmark
Budget: $30 – $60 USD
I'm seeking a proficient python and genAI specialist to wwrite a script to benchmark prompt responses from manuall labelling, response from custom prompt, GPT4, GPT3.5, Mistral using F1-score, precision and recall.
The input will be :
1. pdf document
2. set of questions
3. set of manually extracted answers for the questions
4. set of responses from custom prompt
logic to:
create prompt for GPT4, GPT3.5, Mistral to ask question that are provided in the input from the document
Calculate F1-score, precision and recall.
output
Write questions, answers for manually extracted responses, response from GPT4, GPT3.5, Mistralfor F1-score, precision and recall in an excel or csv file
The input will be :
1. pdf document
2. set of questions
3. set of manually extracted answers for the questions
4. set of responses from custom prompt
logic to:
create prompt for GPT4, GPT3.5, Mistral to ask question that are provided in the input from the document
Calculate F1-score, precision and recall.
output
Write questions, answers for manually extracted responses, response from GPT4, GPT3.5, Mistralfor F1-score, precision and recall in an excel or csv file