Multimodal AI Evaluation Task Creation

Job ID: 40621735

Budget: $8 – $15 USD

Multimodal AI Task Designer

AI Document-Generation Task Author

Overview
Design and build evaluation tasks that test large language models' ability to synthesize multi-source, visually-driven business and technical documents into polished professional deliverables (DOCX, PPTX, XLSX, or PDF). Each task is a full pipeline: sourcing, prompt writing, expert-authored reference output, grading criteria, and model testing, aimed at surfacing specific failure patterns in current AI systems.

Key Responsibilities

Source and curate 4+ real-world documents per task (reports, spreadsheets, slide decks, PDFs, engineering or spec sheets), including at least one distractor file designed to look useful but mislead the model, and at least one file where critical information exists only in a visual element, not extractable as text
Write realistic, paragraph-style workplace prompts (no headers or rigid templates) that establish who is asking, why, what needs to be created, and which sources to use, without prescribing the answer's structure
Author a Golden Response by hand (no AI assistance) that represents the correct, fully realized deliverable at the exact scope the prompt requires
Write structured Grading Guidance covering ground truths, acceptable variation, penalizations, known failure modes, and aesthetic expectations
Generate and refine an objective, binary, unstacked grading rubric from the Grading Guidance, including source citations and justifications for each criterion
Run iterative model evaluations (Gemini 3.5 Flash) and tune task difficulty until the model reliably scores below a 45% average across three runs (all under 55%), confirming the task exposes a genuine model weakness
Pass all required quality gates (Oracle run at 100%, Prompt/Criteria Alignment, Selected Run QC, Human QC) before submitting for Peer Review
Revise tasks in response to Peer Review and HDM feedback through to final sign-off

Skills Required

Strong subject-matter expertise in one or more domains (Business/Consulting, Logistics, Engineering, Architecture, Front-End SWE, or STEM)
Excellent technical writing and prompt design skills
Sharp attention to detail for cross-referencing data across multiple source documents
Working knowledge of DOCX/PPTX/XLSX/PDF authoring and formatting
Understanding of how LLMs typically fail at multimodal and multi-document synthesis tasks, so you can design tasks that expose those gaps

Compensation

Paid per task upon sign-off, at a fixed rate equivalent to 6 hours of work
Reviewer duties compensated at 1 hour equivalent
Paid twice monthly