Website Content Scraper and Spreadsheet Filling via Make.com
Budget: $250 – $750 USD
Project Title: Automate Content Extraction from Website to Spreadsheet Using Make.com
Project Description:
I need a freelancer to set up an automated workflow using Make.com (formerly Integromat) to extract and aggregate content from my website, including blogs, podcasts, YouTube videos, books, and articles. The goal is to populate a structured spreadsheet (Excel or Google Sheets) with this data, making it easy for AI tools to analyze themes, trends, summaries, etc.
Background on the Process:
This is essentially a data extraction or content aggregation task (also known as web scraping or ETL: Extract, Transform, Load). It involves systematically collecting unstructured content from my site and organizing it into a spreadsheet. Each entry should include metadata like title, URL, summaries, tags, and more. I have a template spreadsheet ("new come and reason developer.xlsx") with columns such as:
id
type (e.g., blog, podcast, video, book, article)
title
date_published
source_name
source_url
drive_file_id (if applicable)
text_link
word_count
language
source_topic
primary_theme
secondary_themes
summary_short
summary_long
tags
people_mentioned
status
notes_for_me
You can use this template as the starting point—I'll provide the file upon hiring.
Sources of Content:
Blogs and articles: From my website (I'll provide the URL, RSS feed if available, or sitemap).
Podcasts: Episode links, possibly via RSS; need audio transcription to text.
YouTube videos: From my channel or playlists; fetch metadata and transcribe videos.
Books: Digital excerpts or PDFs; extract text if needed.
Articles: Similar to blogs, from site pages.
Assume a moderate volume initially (e.g., 50-200 items), but the workflow should be scalable for more.
Requirements:
Use Make.com to build the automation scenario(s). Leverage existing templates where possible (e.g., RSS to Google Sheets, web scraping, YouTube integration, AI summarization).
Extract: Pull content via RSS, HTTP requests, or scraping (respect robots.txt).
Transform: Clean text, generate short/long summaries and tags/themes using AI (integrate OpenAI or similar via API), calculate word count, detect language, etc.
Load: Append data row-by-row to the spreadsheet, matching the columns above.
Handle transcription for podcasts/videos: Integrate with tools like AssemblyAI, Otter.ai, or OpenAI Whisper (API costs can be reimbursed or use free tiers for testing).
For books/articles in PDF: Use PDF extraction tools in Make.com.
Make it runnable manually for historical data and schedulable for new content.
Test on a small sample (e.g., 10 items) before full run.
Deliverables:
Fully set up Make.com scenario (share access or export blueprint).
Populated spreadsheet with all current content.
Documentation: Step-by-step guide on how to run/maintain it.
Any custom code if needed (e.g., for complex parsing).
Preferred Skills:
Expertise in Make.com or similar no-code tools (Zapier, Integromat).
Web scraping (e.g., using HTTP modules, parsers).
API integrations (YouTube, OpenAI, transcription services).
Data processing (cleaning, summarization).
Basic Python if custom modules are required (optional).
Budget and Timeline:
Budget: $300-800 (depending on content volume and complexity; please quote based on your assessment).
Timeline: 3-7 days for setup and initial population.
Milestones: 1) Workflow build and test on sample data (50%). 2) Full data population and docs (50%).
Please include in your bid:
Your experience with Make.com and similar projects.
Estimated time/effort.
Any questions about my site/content (I'll share details privately).
I'm looking for someone reliable who can adapt existing templates to save time. If you have pre-built workflows for this, mention them!
Project Description:
I need a freelancer to set up an automated workflow using Make.com (formerly Integromat) to extract and aggregate content from my website, including blogs, podcasts, YouTube videos, books, and articles. The goal is to populate a structured spreadsheet (Excel or Google Sheets) with this data, making it easy for AI tools to analyze themes, trends, summaries, etc.
Background on the Process:
This is essentially a data extraction or content aggregation task (also known as web scraping or ETL: Extract, Transform, Load). It involves systematically collecting unstructured content from my site and organizing it into a spreadsheet. Each entry should include metadata like title, URL, summaries, tags, and more. I have a template spreadsheet ("new come and reason developer.xlsx") with columns such as:
id
type (e.g., blog, podcast, video, book, article)
title
date_published
source_name
source_url
drive_file_id (if applicable)
text_link
word_count
language
source_topic
primary_theme
secondary_themes
summary_short
summary_long
tags
people_mentioned
status
notes_for_me
You can use this template as the starting point—I'll provide the file upon hiring.
Sources of Content:
Blogs and articles: From my website (I'll provide the URL, RSS feed if available, or sitemap).
Podcasts: Episode links, possibly via RSS; need audio transcription to text.
YouTube videos: From my channel or playlists; fetch metadata and transcribe videos.
Books: Digital excerpts or PDFs; extract text if needed.
Articles: Similar to blogs, from site pages.
Assume a moderate volume initially (e.g., 50-200 items), but the workflow should be scalable for more.
Requirements:
Use Make.com to build the automation scenario(s). Leverage existing templates where possible (e.g., RSS to Google Sheets, web scraping, YouTube integration, AI summarization).
Extract: Pull content via RSS, HTTP requests, or scraping (respect robots.txt).
Transform: Clean text, generate short/long summaries and tags/themes using AI (integrate OpenAI or similar via API), calculate word count, detect language, etc.
Load: Append data row-by-row to the spreadsheet, matching the columns above.
Handle transcription for podcasts/videos: Integrate with tools like AssemblyAI, Otter.ai, or OpenAI Whisper (API costs can be reimbursed or use free tiers for testing).
For books/articles in PDF: Use PDF extraction tools in Make.com.
Make it runnable manually for historical data and schedulable for new content.
Test on a small sample (e.g., 10 items) before full run.
Deliverables:
Fully set up Make.com scenario (share access or export blueprint).
Populated spreadsheet with all current content.
Documentation: Step-by-step guide on how to run/maintain it.
Any custom code if needed (e.g., for complex parsing).
Preferred Skills:
Expertise in Make.com or similar no-code tools (Zapier, Integromat).
Web scraping (e.g., using HTTP modules, parsers).
API integrations (YouTube, OpenAI, transcription services).
Data processing (cleaning, summarization).
Basic Python if custom modules are required (optional).
Budget and Timeline:
Budget: $300-800 (depending on content volume and complexity; please quote based on your assessment).
Timeline: 3-7 days for setup and initial population.
Milestones: 1) Workflow build and test on sample data (50%). 2) Full data population and docs (50%).
Please include in your bid:
Your experience with Make.com and similar projects.
Estimated time/effort.
Any questions about my site/content (I'll share details privately).
I'm looking for someone reliable who can adapt existing templates to save time. If you have pre-built workflows for this, mention them!
Related categories:
Data Entry
Excel
Web Scraping
Data Mining
ETL
Automation
API Integration
Make.com