HIRING: Top-Tier AI Data Annotation & Evaluation Specialists -- 3

Job ID: 40505946

Budget: $2 – $8 USD

a) Setting a concrete Rubric-based evaluation criteria (as is formally done in LLMs' response evaluation) such as:
- Accuracy
- Instruction following
- Reasoning quality
- Completeness
- Clarity
- Safety
- Coding correctness
- Hallucination risk

b) To help optimize prompts, prompt engineering frameworks can be employed such as zero-shot prompting, few-shot prompting, chain-of-thought prompting.

c) All errors in the model's responses can be segregated into different error types like Factual error, Reasoning error, Instruction failure, Formatting issue, Hallucination, Safety issue, Code bug, etc so that we can deeply analyse the performance so all the models and figure out their specific weaknesses.