AI Evaluation: Identifying ChatGPT Failures

Job ID: 40530135

Budget: $10 – $30 USD

You will create realistic search prompts that cause ChatGPT to make a verifiable search or retrieval error. The goal is to identify cases where the model provides incorrect information, uses outdated sources, relies on low-quality websites, or fails to find information available on authoritative sources.

What You'll Do

Create objective, real-world prompts with a single verifiable answer.
Test prompts in ChatGPT and identify valid failures.
Verify the correct answer using reliable sources.
Document what went wrong in the model's response.
Categorize the failure type (e.g., outdated information, wrong source, niche source not found, spam/SEO source).
Write clear instructions that lead to the correct answer.
Create a rubric to evaluate future model responses.

Requirements

Strong research and fact-checking skills.
Excellent attention to detail.
Ability to evaluate source quality and credibility.
Clear written English.
Familiarity with ChatGPT or other AI tools.

Success Criteria
Find genuine, reproducible AI search failures and provide clear documentation, validation steps, and evaluation rubrics that can be used to improve model performance.