Job Title: Senior AI Evaluation Specialist (Coding & Codebase Analysis) $3hr

Job ID: 40476237

Budget: $2 – $8 USD

Location: RemoteType: Contract / FreelanceAbout the RoleWe are seeking a highly meticulous, tech-savvy Senior AI Evaluation Specialist to audit and grade AI-generated code solutions. In this role, you will act as the ultimate quality assurance judge. You will review complex coding transcripts, analyze how different AI models interact with repositories, inspect final code outputs for absolute correctness, and write data-driven, evidence-based evaluation reports.This is not a traditional software engineering role. You do not need to write code from scratch; instead, you need the elite code-reading architecture to spot subtle edge-case failures, evaluate tool-use patterns, and judge code quality purely based on project transcripts.Key ResponsibilitiesComprehensive Code Auditing: Deeply analyze two competing AI model transcripts (Response A and Response B) responding to complex coding prompts.Outcome-Focused Grading: Evaluate final code outputs for technical correctness, safety, efficiency, and architectural integrity.Taxonomy-Driven Error Mapping: Identify and log precise behavioral weaknesses using a structured taxonomy (e.g., Instruction Following Failures, Overengineering, Tool Use Errors, Laziness).Technical Writing & Justification: Write detailed, objective, and evidence-based score rationales, citing exact file names, tool calls, and lines of code.Required QualificationsStrong Code Literacy: Ability to easily read, interpret, and mentally trace code across multiple languages and modern web/data frameworks.Elite Attention to Detail: Experience adhering to strict evaluation rubrics, exact character counts, and complex formatting guidelines.Analytical Writing Skills: Proven capability to write structured, objective summaries in plain, professional English, translating technical errors into clear rationales.Skepticism & Objectivity: A mindset that prioritizes final code correctness over a "clean" process, ensuring no bias slips into the final grading.