Create Labelled Data and the LLM Image Model
Budget: $750 – $1,500 USD
I need a complete, end-to-end solution that pairs a rigorously labeled image dataset with a production-ready model capable of reaching a mean absolute error of 2.5 % or lower against my own ratings, which follow a detailed Rating Reference Guide I will supply.
Full rating guide inside, check it for all details: https://docs.google.com/document/d/1bxnvOu4lkvoxSTliYyUZfNaKA8GYodkxgFM2IlBrAao/edit?usp=sharing
Summary Scope
• Curate or collect images, then label them precisely according to the guide. Data quality, balance, and clear provenance are essential.
• Design and train a model—architecture is flexible—but it must generalise well, remain lightweight enough for real-time inference, and achieve ≤ 2.5 % MAE on an unseen hold-out set I provide.
• Document preprocessing, augmentation, and hyper-parameter choices so I can reproduce results.
• Package everything for production: cleaned dataset, training scripts, inference pipeline, and a quick-start README. A containerised build (e.g., Docker) is preferred for friction-free deployment.
• Supply validation reports: learning curves, confusion analyses, error heat-maps, and a concise write-up of limitations plus recommended next steps to harden robustness.
Summary Acceptance criteria
1. Labeled dataset passes a manual spot-check for consistency and completeness.
2. Model achieves the target MAE when run against my blind test set on identical hardware.
3. Codebase is version-controlled, well-commented, and runnable with a single command.
4. Final hand-off includes the trained weights, scripts, dataset in agreed format, and documentation.
Full rating guide inside, check it for all details: https://docs.google.com/document/d/1bxnvOu4lkvoxSTliYyUZfNaKA8GYodkxgFM2IlBrAao/edit?usp=sharing
Summary Scope
• Curate or collect images, then label them precisely according to the guide. Data quality, balance, and clear provenance are essential.
• Design and train a model—architecture is flexible—but it must generalise well, remain lightweight enough for real-time inference, and achieve ≤ 2.5 % MAE on an unseen hold-out set I provide.
• Document preprocessing, augmentation, and hyper-parameter choices so I can reproduce results.
• Package everything for production: cleaned dataset, training scripts, inference pipeline, and a quick-start README. A containerised build (e.g., Docker) is preferred for friction-free deployment.
• Supply validation reports: learning curves, confusion analyses, error heat-maps, and a concise write-up of limitations plus recommended next steps to harden robustness.
Summary Acceptance criteria
1. Labeled dataset passes a manual spot-check for consistency and completeness.
2. Model achieves the target MAE when run against my blind test set on identical hardware.
3. Codebase is version-controlled, well-commented, and runnable with a single command.
4. Final hand-off includes the trained weights, scripts, dataset in agreed format, and documentation.