OpenClaw RL / Outlier AI Evaluation Engineer
Budget: $8 – $15 USD
Overview:
We are looking for an AI evaluation specialist with experience in OpenClaw, Outlier, LLM evaluation, Python, and rubric design. The role involves designing complex OpenClaw agent tasks, evaluating model trajectories, writing rubrics, creating Python pytest verifiers, and identifying safety or quality failures.
Responsibilities:
- Design complex multi-stage OpenClaw agent tasks
- Create natural prompts with clear objectives and final artifacts
- Run or evaluate model trajectories across different LLMs
- Extract and review OpenClaw trajectories and workspace outputs
- Build objective, atomic, self-contained rubrics
- Write Python pytest unit tests in verifiers.py
- Evaluate tool usage, reasoning, instruction following, and final output quality
- Identify and annotate safety failures using project taxonomy
- Compare, rate, and rank model performance
- Follow all platform guidelines, privacy rules, and data compliance requirements
Requirements:
- Experience with OpenClaw, Outlier, or similar AI evaluation platforms
- Strong understanding of LLM agents and tool-use workflows
- Python experience, especially pytest/unit testing
- Ability to write clear rubrics and evaluation criteria
- Strong English reading and writing skills
- Detail-oriented and able to follow complex project instructions
- Understanding of AI safety, privacy, and prompt injection risks
- Ability to work remotely and asynchronously
Nice to Have:
- Prior OpenClaw Atlas or OpenClaw RL project experience
- Experience with RLHF, model ranking, or trajectory evaluation
- Familiarity with browser tools, APIs, spreadsheets, email/calendar workflows
- Experience reviewing safety failures or AI task quality
We are looking for an AI evaluation specialist with experience in OpenClaw, Outlier, LLM evaluation, Python, and rubric design. The role involves designing complex OpenClaw agent tasks, evaluating model trajectories, writing rubrics, creating Python pytest verifiers, and identifying safety or quality failures.
Responsibilities:
- Design complex multi-stage OpenClaw agent tasks
- Create natural prompts with clear objectives and final artifacts
- Run or evaluate model trajectories across different LLMs
- Extract and review OpenClaw trajectories and workspace outputs
- Build objective, atomic, self-contained rubrics
- Write Python pytest unit tests in verifiers.py
- Evaluate tool usage, reasoning, instruction following, and final output quality
- Identify and annotate safety failures using project taxonomy
- Compare, rate, and rank model performance
- Follow all platform guidelines, privacy rules, and data compliance requirements
Requirements:
- Experience with OpenClaw, Outlier, or similar AI evaluation platforms
- Strong understanding of LLM agents and tool-use workflows
- Python experience, especially pytest/unit testing
- Ability to write clear rubrics and evaluation criteria
- Strong English reading and writing skills
- Detail-oriented and able to follow complex project instructions
- Understanding of AI safety, privacy, and prompt injection risks
- Ability to work remotely and asynchronously
Nice to Have:
- Prior OpenClaw Atlas or OpenClaw RL project experience
- Experience with RLHF, model ranking, or trajectory evaluation
- Familiarity with browser tools, APIs, spreadsheets, email/calendar workflows
- Experience reviewing safety failures or AI task quality