ChatGPT Output Evaluation Pipeline
Budget: $250 – $750 USD
I'm looking for a freelancer to build an evaluation pipeline that can test the output quality of the ChatGPT web app (browser-based, not API).
The pipeline should simulate user interactions through the interface and measure:
- Response relevance and helpfulness
- Latency (time to first token and full response)
- Consistency across repeated prompts
Requirements:
- Experience with browser automation (e.g., Playwright, Puppeteer, or Selenium)
- Familiarity with evaluating AI-generated outputs
- Ability to log interactions and structure outputs for analysis
Budget: $500 — This is for the first phase of the project. We have a larger roadmap ahead and are looking for someone reliable. If you deliver well, we’re happy to continue working together and increase the budget for follow-on tasks.
Please include relevant experience or similar work in your bid.
The pipeline should simulate user interactions through the interface and measure:
- Response relevance and helpfulness
- Latency (time to first token and full response)
- Consistency across repeated prompts
Requirements:
- Experience with browser automation (e.g., Playwright, Puppeteer, or Selenium)
- Familiarity with evaluating AI-generated outputs
- Ability to log interactions and structure outputs for analysis
Budget: $500 — This is for the first phase of the project. We have a larger roadmap ahead and are looking for someone reliable. If you deliver well, we’re happy to continue working together and increase the budget for follow-on tasks.
Please include relevant experience or similar work in your bid.