Curate and Script Large-Scale AI Dataset Downloads (Tech-Focused)
Budget: $250 – $750 USD
We’re seeking a skilled, self-sufficient data expert to assist in downloading and managing a large set of AI training datasets, beginning with the technical category.
This project requires:
• Independent link collection: You must locate and gather the dataset links yourself. Resourcefulness is key—no spoon-feeding.
• Phase One: Technical data — must be completed within 7 days of hire.
• Download scripting expertise:
○ Must throttle downloads (approx. 350–400GB/day to avoid ISP caps)
○ Must support resume after interruption
○ Must ensure file integrity
• No downloads from Dropbox, Google Drive, OneDrive, or similar “user storage” sites. Links should point to direct-access datasets hosted on institutional or research-grade sources (Hugging Face, S3, GitHub, etc.).
• Clear documentation of all steps, links, and scripts used
Ideal Candidate:
• Experienced with high-volume dataset handling
• Comfortable with CLI tools like wget, aria2, rsync, or similar
• Understands structured, efficient downloads across large-scale file systems
• Familiar with Wasabi, S3, or equivalent cloud storage systems
Project Scope:
• This is the first stage (technical data) in a multi-part collection
• Full project spans multiple categories including legal, medical, and narrative datasets
• This is a mission-critical private AI project—clean execution, secure handling, and reliability are a must
Apply only if you’re confident in working with massive datasets and can deliver with speed and precision, without oversight.
This project requires:
• Independent link collection: You must locate and gather the dataset links yourself. Resourcefulness is key—no spoon-feeding.
• Phase One: Technical data — must be completed within 7 days of hire.
• Download scripting expertise:
○ Must throttle downloads (approx. 350–400GB/day to avoid ISP caps)
○ Must support resume after interruption
○ Must ensure file integrity
• No downloads from Dropbox, Google Drive, OneDrive, or similar “user storage” sites. Links should point to direct-access datasets hosted on institutional or research-grade sources (Hugging Face, S3, GitHub, etc.).
• Clear documentation of all steps, links, and scripts used
Ideal Candidate:
• Experienced with high-volume dataset handling
• Comfortable with CLI tools like wget, aria2, rsync, or similar
• Understands structured, efficient downloads across large-scale file systems
• Familiar with Wasabi, S3, or equivalent cloud storage systems
Project Scope:
• This is the first stage (technical data) in a multi-part collection
• Full project spans multiple categories including legal, medical, and narrative datasets
• This is a mission-critical private AI project—clean execution, secure handling, and reliability are a must
Apply only if you’re confident in working with massive datasets and can deliver with speed and precision, without oversight.