Training Data Collection
Budget: $10 – $70 USD
Objective:
We are seeking a data collector to gather structured textual data for training a bot. The goal is to create a dataset comprising high-quality content in TXT or JSON format, suitable for training a model to understand humor, interviews, and social media interactions.
Scope of Work
Stand-Up Transcripts:
Collect the full text transcripts of every stand-up comedy special by Dave Chappelle and Bill Burr.
Structure the data clearly by including:
Title of the special
Year of release
Segmented jokes or monologues
Context (if any, like audience interaction)
Interviews:
Gather the textual content from 100 interviews of Dave Chappelle and Bill Burr combined.
Ensure the interviews span different years and topics to capture a variety of perspectives.
Include metadata:
Interviewee (Chappelle or Burr)
Interviewer
Source (e.g., podcast, TV, magazine)
Date
Social Media Content:
Extract all tweets from 5 tweeters like:
crypto_bitlord7
idrawline
Include:
Tweet text
We are seeking a data collector to gather structured textual data for training a bot. The goal is to create a dataset comprising high-quality content in TXT or JSON format, suitable for training a model to understand humor, interviews, and social media interactions.
Scope of Work
Stand-Up Transcripts:
Collect the full text transcripts of every stand-up comedy special by Dave Chappelle and Bill Burr.
Structure the data clearly by including:
Title of the special
Year of release
Segmented jokes or monologues
Context (if any, like audience interaction)
Interviews:
Gather the textual content from 100 interviews of Dave Chappelle and Bill Burr combined.
Ensure the interviews span different years and topics to capture a variety of perspectives.
Include metadata:
Interviewee (Chappelle or Burr)
Interviewer
Source (e.g., podcast, TV, magazine)
Date
Social Media Content:
Extract all tweets from 5 tweeters like:
crypto_bitlord7
idrawline
Include:
Tweet text