Training with Webscraped Data for Attractiveness LLM

Job ID: 39748657

Budget: $250 – $750 USD

Please Read this:
https://drive.google.com/file/d/1oVEI_29BRd1Kwh09s58MztAC8_L_Csoh/view?usp=sharing

This project involves web scraping a specified website to collect data on a single property—attractiveness—and using that data to train a Large Language Model (LLM) via on-policy distillation. The goal is to create a distilled model that replicates the website's output for attractiveness assessment with high accuracy. This is a small-scale project, and the implementer expert is expected to leverage your existing LLM training setup, so please allocate resources in your quote. If this takes longer than a week, please make sure to provide daily updates on the delay. Successful completion may lead to additional similar projects, where resources can be provided by the client.

The entire solution must be encapsulated in a Docker container. The container should run on a preassigned port and utilize cloud resources (e.g., for testing and inference). Free methods and tools are encouraged wherever possible. The LLM must be trained specifically on data from the provided website (the simple website just has an image upload on the first page - easy to scrape and on-policy).