Ongoing College Soccer Data Scraper, and Database Manager

Job ID: 40184094

Budget: $250 – $750 USD

I run an app that guides high-school soccer players through the college-recruiting maze, and the quality of our data is what keeps families coming back. I’m ready to hand over the scraping and database upkeep to someone who can treat it as a season-long partnership rather than a one-off gig.

The work centers on three source groups: mainstream college sports sites, broader college-data portals, and the third-party pages that promote soccer ID camps. From those locations I need player statistics, coach and staff contact details, plus dates, costs, and locations for every camp or showcase we list. Much of the information refreshes monthly—sometimes weekly when schedules change—while other datasets only need a deep clean once or twice a year. Because coaches can switch schools overnight, tight turnaround on last-minute updates is critical, so I need you to be based in the United States, Canada (or the EU if the right person) for quick, daytime communication. I WILL ONLY HIRE SOMEONE WHOS PROFILE PICTURE MATCHES WHO THEY ACTUALLY ARE. A VIDEO CALL TO DISCUSS WILL BE NEEDED TO DISCUSS DETAILS PRIOR TO HIRING.

**THIS IS A LONG TERM PROJECT THAT WILL BE HEAVILY FRONT LOADED HOPEFULLY MAKING YOUR LIFE EASIER OVER THE DURATION OF THE PROJECT. WITH THIS I PLAN ON PROVIDING A MONTHLY STIPEND THAT WILL EVEN OUT PAY OVER THAT TERM.

Your toolkit is up to you, but Python with Scrapy or Selenium, a tidy Pandas workflow, and solid MySQL / PostgreSQL skills fit naturally with what we already run on AWS. Clean code, deduping, error logging, and clear documentation are non-negotiable; everything has to slip straight into the pipelines that feed our mobile app.

** I HAVE EXISTING SCRIPTS FROM THE PREVIOUS FREELANCER THAT WILL HELP YOU EXPEDITE THE WORK YOU DO.

Deliverables
• Initial full scrape of all listed sites, normalized into the existing schema
• Scheduled or on-demand scripts for incremental updates, ready to run via cron or Lambda
• A simple dashboard or report that flags changed rows and any scrape failures
• Documentation good enough that another developer could step in without guessing

I’ll map out the calendar of expected refreshes once we’ve agreed on the workflow, but the sooner the first complete dataset lands, the sooner our athletes get the insight they need. Let’s set up a quick call and see if our timelines and tech stacks match up.

If interested I can immediately send you the first tabulated spreadsheets so you can get a scope of one of the projects.

Also