Python Specialist for AI Agent Evaluation Framework

Job ID: 40573384

Budget: $250 – $750 USD

**Title:** Python Developer — Build Open-Source AI Agent Evaluation Framework (Supply Chain)

**About the project:**
Building an open-source evaluation framework that tests how well AI agents make supply chain planning decisions. Final output will be a public GitHub repository. Not a commercial product — a research artifact.

**What you will build:**
- A synthetic supply chain dataset generator (network, vendors, BOM, demand)
- 2-3 decision problem families (demand planning, replenishment, procurement)
- A reference solver in Pyomo + HiGHS to compute optimal answers
- An evaluation harness that scores any AI agent against the reference
- 3 baseline agents: one LLM-based (LangGraph + Claude/GPT), one optimization-only, one heuristic
- A scoring module measuring quality, reliability across repeated runs, and efficiency

**Tools (all open source):**
Python, LangGraph, LangChain, Pyomo, HiGHS, Pandas, NumPy, Pytest, GitHub.

**What I provide:**
Domain framing, decision family specs, math formulations for the reference solver, scoring methodology, and ongoing technical direction.

**Requirements:**
- Strong Python (5+ years)
- LangGraph or LangChain experience (share GitHub links)
- Pyomo or similar optimization framework
- Prior open-source contributions
- Clean code with tests and documentation

**Nice to have:**
- Supply chain or operations research background
- Prior work on evaluation frameworks or benchmarks

**Engagement:**
- Freelance, milestone-based
- Weekly 30-minute video check-ins

**How to apply:**
Share your hourly rate or fixed bid, GitHub or portfolio links, and any relevant code samples