ChatGPT Output Evaluation Pipeline

Job ID: 39539676

Budget: $250 – $750 USD

I'm looking for a freelancer to build an evaluation pipeline that can test the output quality of the ChatGPT web app (browser-based, not API).

The pipeline should simulate user interactions through the interface and measure:
- Response relevance and helpfulness
- Latency (time to first token and full response)
- Consistency across repeated prompts

Requirements:
- Experience with browser automation (e.g., Playwright, Puppeteer, or Selenium)
- Familiarity with evaluating AI-generated outputs
- Ability to log interactions and structure outputs for analysis

Budget: $500 — This is for the first phase of the project. We have a larger roadmap ahead and are looking for someone reliable. If you deliver well, we’re happy to continue working together and increase the budget for follow-on tasks.

Please include relevant experience or similar work in your bid.