AI Function Calling Evaluation Specialist for Insurance Claim Assistant

Job ID: 40435503

Budget: $30 – $250 USD

We are looking for detail-oriented AI evaluators to test and label conversations with an insurance claim assistant AI model. This project focuses on checking whether the AI correctly triggers function calls during realistic claim-related conversations, such as verifying policies, creating claims, assessing damage, checking claim status, adding notes, flagging cases for adjusters, and scheduling inspections.

The work involves recording natural voice conversations with the AI, using provided fake user and policy information, and evaluating whether the model responds correctly. You will need to include multiple function-call scenarios, show a damage image for assessment, test behavior under background noise, and complete required quality checks such as long-number readback, short replies, pronunciation, emotional support, safety refusal, and system/JSON leakage behavior.

Responsibilities:

Conduct 4 to 8 minute recorded conversations with an AI claim assistant.

Use only provided fake customer and policy information.

Create realistic insurance claim scenarios involving auto, property, water, fire, theft, or similar claim types.

Trigger and evaluate function calls such as policy verification, claim creation, damage assessment, adjuster flagging, claim status checks, claim notes, and inspection scheduling.

Upload or show a clear damage image during the conversation and evaluate whether the AI handles damage assessment correctly.

Complete classification questions after each recording.

Write ideal function calls in the required Python-style format with accurate timestamps and parameters.

Identify model errors, missed function calls, wrong parameters, poor responses, voice issues, interruptions, safety failures, or system leakage issues.

Revise rejected tasks when needed based on reviewer feedback.

Requirements:

Strong English communication skills.

Ability to speak naturally and clearly during recorded AI conversations.

Good attention to detail.

Comfortable following detailed project instructions.

Basic understanding of AI tools, function calling, or chatbot evaluation is preferred.

Ability to write structured notes and follow exact formatting rules.

Reliable internet connection and clear microphone/audio setup.

Must be able to create or use background noise when required by the task instructions.

Must not use personal information or real customer information in conversations.

Ideal Candidate:

You are careful, patient, and good at following instructions. You can role-play naturally, notice small mistakes in AI behavior, and clearly explain what should have happened. Prior experience with AI data labeling, QA testing, chatbot testing, voice AI evaluation, or insurance-related workflows is a plus.