Firebase Article SubscriptiExpert Python Developer for Newspaper Scraping & OCR: Google Drive & Firebase Integrationon Web App
Budget: ₹1,500 – ₹12,500 INR
I am looking for an experienced Python Developer to automate a data pipeline for my Public Notice Aggregator website. I have a curated list of 500+ E-paper URLs. The goal is to extract public notices daily and store them systematically.
The Workflow:
Automated Scraping: Daily extraction of 'Public Notice' images/sections from 500+ e-paper websites based on the provided list.
Storage: All extracted images must be uploaded to my Google Drive (via Google Drive API) into organized folders (Date/District wise). I do not want to store images on Firebase.
OCR & Parsing: Use Google Cloud Vision API to extract text from these images (supporting Marathi and Hindi languages).
Data Formatting: Every notice text must start with a mandatory source line.
Format Example: स्त्रोत : [Newspaper Name], [Edition], प्रकाशन दिनांक [Date]
Database Update: The final formatted text (with the source line) must be pushed to my Firebase Firestore database.
Cost Optimization: The entire script should be optimized to run on Google Cloud Functions to keep operational costs at a minimum.
Technical Requirements:
Proficiency in Python and web scraping libraries (Selenium, Scrapy, etc.).
Experience with Google Drive API and Google Cloud Vision API.
Strong understanding of Firebase Firestore and Cloud Functions.
Ability to handle Marathi/Hindi OCR accuracy and font formatting.
Experience using Gemini API for automated text cleaning and formatting is a plus.
Deliverables:
Fully functional automation script.
Integration with my existing Firebase/React setup.
Clear documentation on how to update the newspaper URL list in the future.
The Workflow:
Automated Scraping: Daily extraction of 'Public Notice' images/sections from 500+ e-paper websites based on the provided list.
Storage: All extracted images must be uploaded to my Google Drive (via Google Drive API) into organized folders (Date/District wise). I do not want to store images on Firebase.
OCR & Parsing: Use Google Cloud Vision API to extract text from these images (supporting Marathi and Hindi languages).
Data Formatting: Every notice text must start with a mandatory source line.
Format Example: स्त्रोत : [Newspaper Name], [Edition], प्रकाशन दिनांक [Date]
Database Update: The final formatted text (with the source line) must be pushed to my Firebase Firestore database.
Cost Optimization: The entire script should be optimized to run on Google Cloud Functions to keep operational costs at a minimum.
Technical Requirements:
Proficiency in Python and web scraping libraries (Selenium, Scrapy, etc.).
Experience with Google Drive API and Google Cloud Vision API.
Strong understanding of Firebase Firestore and Cloud Functions.
Ability to handle Marathi/Hindi OCR accuracy and font formatting.
Experience using Gemini API for automated text cleaning and formatting is a plus.
Deliverables:
Fully functional automation script.
Integration with my existing Firebase/React setup.
Clear documentation on how to update the newspaper URL list in the future.
Related categories:
Python
Web Scraping
OCR
Node.js
Image Processing
Selenium
Automation
Database Management
API Integration