Data Scraping and Database Setup Specialist
Budget: $30 – $250 USD
NEED THIS DONE ASAP! DONT APPLY IF YOU CANT START NOW! AND DONT HAVE EXPERIENCE!
We are seeking a skilled developer to handle data scraping and database setup for a high-performance chatbot project. This role involves collecting structured data from multiple websites and designing a robust database to store and retrieve the information efficiently.
Responsibilities
Data Scraping
Identify target websites and develop web scrapers using tools like BeautifulSoup, Scrapy, or Selenium.
Automate the scraping process using AWS Lambda or other scheduling tools.
Extract relevant data such as schedules, events, and other updates from dynamic or static websites.
Clean and format the scraped data into structured formats (e.g., JSON, CSV).
Database Setup
Design and implement a schema for storing the scraped data in Amazon DynamoDB.
Configure DynamoDB for high-throughput data storage with replication and backup.
Populate the database with scraped data, ensuring optimized indexing for quick retrieval.
Perform test queries to verify the database’s performance and accuracy.
Integration
Collaborate with the backend and chatbot development teams to ensure smooth integration of the database into the overall system.
Periodically update the database with new data through automated scraping processes.
Required Skills
Web Scraping:
Proficiency in Python web scraping libraries like BeautifulSoup, Scrapy, and Selenium.
Experience in handling dynamic websites and APIs.
Database Expertise:
Strong knowledge of Amazon DynamoDB and database design.
Ability to optimize database queries and structure for high performance.
AWS Experience:
Familiarity with AWS Lambda and CloudWatch for automation and monitoring.
Basic understanding of AWS services for data handling and storage.
Data Cleaning and Transformation:
Proficiency in libraries like Pandas for cleaning and transforming scraped data.
Deliverables:
Functional web scrapers for target websites.
A well-structured and populated DynamoDB database.
Documentation on scraping processes and database schema.
Automated scraping scripts scheduled for periodic updates.
We are seeking a skilled developer to handle data scraping and database setup for a high-performance chatbot project. This role involves collecting structured data from multiple websites and designing a robust database to store and retrieve the information efficiently.
Responsibilities
Data Scraping
Identify target websites and develop web scrapers using tools like BeautifulSoup, Scrapy, or Selenium.
Automate the scraping process using AWS Lambda or other scheduling tools.
Extract relevant data such as schedules, events, and other updates from dynamic or static websites.
Clean and format the scraped data into structured formats (e.g., JSON, CSV).
Database Setup
Design and implement a schema for storing the scraped data in Amazon DynamoDB.
Configure DynamoDB for high-throughput data storage with replication and backup.
Populate the database with scraped data, ensuring optimized indexing for quick retrieval.
Perform test queries to verify the database’s performance and accuracy.
Integration
Collaborate with the backend and chatbot development teams to ensure smooth integration of the database into the overall system.
Periodically update the database with new data through automated scraping processes.
Required Skills
Web Scraping:
Proficiency in Python web scraping libraries like BeautifulSoup, Scrapy, and Selenium.
Experience in handling dynamic websites and APIs.
Database Expertise:
Strong knowledge of Amazon DynamoDB and database design.
Ability to optimize database queries and structure for high performance.
AWS Experience:
Familiarity with AWS Lambda and CloudWatch for automation and monitoring.
Basic understanding of AWS services for data handling and storage.
Data Cleaning and Transformation:
Proficiency in libraries like Pandas for cleaning and transforming scraped data.
Deliverables:
Functional web scrapers for target websites.
A well-structured and populated DynamoDB database.
Documentation on scraping processes and database schema.
Automated scraping scripts scheduled for periodic updates.