Design Deterministic Tech Agent Scenarios
Budget: $25 – $50 USD
I have a collection of synthetic YAML “seed” files that describe high-level tasks. Your job is to turn each seed into a fully fleshed-out, deterministic scenario set in the real world—specifically within Technology and, even more narrowly, everyday Software Development work.
Here’s the flow I need you to own from end to end:
• World crafting
Refine the seed so the environment, constraints, and resources feel authentic to professional software engineers. Every step must stay deterministic: given the same inputs, the agent should always encounter the same world state.
• Multi-model run-through
Execute the scenario against at least two distinct agent models. One model must succeed, another must fail. Capture all prompts, intermediate thoughts, and outputs so results can be reproduced.
• Solution outline
Write a concise reasoning path that shows how an ideal agent reaches success. This becomes the canonical answer key.
• QA audit
Compare the automated grading system’s verdicts against your own ground truth, and explain any mismatches so we can tighten the evaluation logic.
Deliverables for each seed:
1. Final YAML scenario file
2. Log bundle of agent runs (pass & fail)
3. Solution outline (step-by-step reasoning)
4. QA audit report with recommendations
If you’re comfortable juggling YAML, agent frameworks (e.g., OpenAI, Claude, or similar), and meticulous reproducibility requirements, I’d love to see how you approach one of these seeds and iterate from there.
Here’s the flow I need you to own from end to end:
• World crafting
Refine the seed so the environment, constraints, and resources feel authentic to professional software engineers. Every step must stay deterministic: given the same inputs, the agent should always encounter the same world state.
• Multi-model run-through
Execute the scenario against at least two distinct agent models. One model must succeed, another must fail. Capture all prompts, intermediate thoughts, and outputs so results can be reproduced.
• Solution outline
Write a concise reasoning path that shows how an ideal agent reaches success. This becomes the canonical answer key.
• QA audit
Compare the automated grading system’s verdicts against your own ground truth, and explain any mismatches so we can tighten the evaluation logic.
Deliverables for each seed:
1. Final YAML scenario file
2. Log bundle of agent runs (pass & fail)
3. Solution outline (step-by-step reasoning)
4. QA audit report with recommendations
If you’re comfortable juggling YAML, agent frameworks (e.g., OpenAI, Claude, or similar), and meticulous reproducibility requirements, I’d love to see how you approach one of these seeds and iterate from there.