AI Response Quality Rater

Job ID: 39932937

Budget: $250 – $750 USD

I need a few detail-oriented raters located in the United States to score and comment on answers generated by large language models. You will log in through our web portal, read each model response, and decide whether it meets quality, relevance, and safety guidelines that I will provide.

Scope of work
• Commit to roughly 5–6 hours of rating per day, Monday through Friday.
• Follow a short online training that explains our rubric and common edge cases.
• Evaluate responses with a focus on quality assessment; you may occasionally flag bias or factual errors when you spot them.
• Enter clear, concise feedback so we can refine future model updates.

What I’m looking for
• Some prior experience using or evaluating AI/LLM systems—if you have worked with BERT-based models or similar, that is a plus.
• Consistent internet connection and your own laptop or desktop.
• Ability to stay objective while reading a high volume of text.

Payment
The role is paid at $16 per hour, issued weekly through Freelancer’s tracker.

If this sounds like a fit, send a brief note outlining your AI experience and the earliest date you could start. I’ll reach out with next steps and onboarding details.