Curate and Script Large-Scale AI Dataset Downloads (Tech-Focused)

Job ID: 39611129

Budget: $250 – $750 USD

We’re seeking a skilled, self-sufficient data expert to assist in downloading and managing a large set of AI training datasets, beginning with the technical category.

This project requires:

• Independent link collection: You must locate and gather the dataset links yourself. Resourcefulness is key—no spoon-feeding.
• Phase One: Technical data — must be completed within 7 days of hire.
• Download scripting expertise:
○ Must throttle downloads (approx. 350–400GB/day to avoid ISP caps)
○ Must support resume after interruption
○ Must ensure file integrity
• No downloads from Dropbox, Google Drive, OneDrive, or similar “user storage” sites. Links should point to direct-access datasets hosted on institutional or research-grade sources (Hugging Face, S3, GitHub, etc.).
• Clear documentation of all steps, links, and scripts used

Ideal Candidate:

• Experienced with high-volume dataset handling
• Comfortable with CLI tools like wget, aria2, rsync, or similar
• Understands structured, efficient downloads across large-scale file systems
• Familiar with Wasabi, S3, or equivalent cloud storage systems

Project Scope:

• This is the first stage (technical data) in a multi-part collection
• Full project spans multiple categories including legal, medical, and narrative datasets
• This is a mission-critical private AI project—clean execution, secure handling, and reliability are a must

Apply only if you’re confident in working with massive datasets and can deliver with speed and precision, without oversight.