HIRING: Top-Tier AI Data Annotation & Evaluation Specialists -- 7
Budget: $2 – $8 USD
AI Agent Evaluator
Remote · Contract · Flexible Hours
We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained.
What You'll Do
You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project.
You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output.
You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set.
You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier.
You will write pytest-based unit tests in a verifiers.py file that validate the agent's final system state using snapshots.json as ground truth — confirming outcomes, not just intent.
You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace.
You Must Be Able To
Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes
Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models
Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier
Read and reason critically about full agent trajectories
Write and validate basic Python unit tests using pytest
Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios
Strong Background In
AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking.
Important
This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.
Remote · Contract · Flexible Hours
We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained.
What You'll Do
You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project.
You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output.
You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set.
You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier.
You will write pytest-based unit tests in a verifiers.py file that validate the agent's final system state using snapshots.json as ground truth — confirming outcomes, not just intent.
You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace.
You Must Be Able To
Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes
Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models
Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier
Read and reason critically about full agent trajectories
Write and validate basic Python unit tests using pytest
Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios
Strong Background In
AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking.
Important
This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.