Backend Development for Intelligent Monitoring of Government Websites and Document Analysis
Budget: $30 – $250 USD
Summary
We are seeking a backend developer to create a robust and flexible system for monitoring public agency websites. The main objective is to identify and record occurrences of specific keywords within published documents, with a special focus on PDF files.
System Objectives:
* Automated access to various public agency websites.
* Identification and download of newly published documents.
* Reading and extraction of internal content from documents.
* Precise location of predefined keywords.
* Detailed recording of each occurrence (data, source, excerpt, and keyword found).
Challenges and Technical Requirements:
* Flexibility and Modularity: Public agency websites do not follow a single technical framework. The system must be able to handle diverse HTML structures, different ways of listing documents, dynamic URLs, and publications in multiple formats (PDF, HTML, etc.), in addition to pagination and filters specific to each website. The architecture must allow the implementation of specific scraping strategies per source.
* Overcoming Protections: Some websites may present CAPTCHAs, blocks due to excessive requests, anti-bot protections, and content loaded via JavaScript. The backend must be able to detect captchas, log access failures, and implement technical fallbacks (such as the use of headless browsers) ethically and legally, without illegally bypassing security systems.
* Document Content Analysis: The critical functionality is keyword searching WITHIN the document content, not just on web pages. This requires the ability to download files, extract internal text (mainly from PDFs), normalize text (handling accents, instructions/lowercase, etc.), and perform highly accurate searches to avoid false positives.
* Keyword Management: The system must offer an interface to register and manage multiple keywords, associating search results with the correct user and recording the relevant excerpt, data, and document source.
We are looking for a professional with strong knowledge in backend development, advanced web scraping, and document processing, capable of building a scalable and high-performance solution for this project.
We are seeking a backend developer to create a robust and flexible system for monitoring public agency websites. The main objective is to identify and record occurrences of specific keywords within published documents, with a special focus on PDF files.
System Objectives:
* Automated access to various public agency websites.
* Identification and download of newly published documents.
* Reading and extraction of internal content from documents.
* Precise location of predefined keywords.
* Detailed recording of each occurrence (data, source, excerpt, and keyword found).
Challenges and Technical Requirements:
* Flexibility and Modularity: Public agency websites do not follow a single technical framework. The system must be able to handle diverse HTML structures, different ways of listing documents, dynamic URLs, and publications in multiple formats (PDF, HTML, etc.), in addition to pagination and filters specific to each website. The architecture must allow the implementation of specific scraping strategies per source.
* Overcoming Protections: Some websites may present CAPTCHAs, blocks due to excessive requests, anti-bot protections, and content loaded via JavaScript. The backend must be able to detect captchas, log access failures, and implement technical fallbacks (such as the use of headless browsers) ethically and legally, without illegally bypassing security systems.
* Document Content Analysis: The critical functionality is keyword searching WITHIN the document content, not just on web pages. This requires the ability to download files, extract internal text (mainly from PDFs), normalize text (handling accents, instructions/lowercase, etc.), and perform highly accurate searches to avoid false positives.
* Keyword Management: The system must offer an interface to register and manage multiple keywords, associating search results with the correct user and recording the relevant excerpt, data, and document source.
We are looking for a professional with strong knowledge in backend development, advanced web scraping, and document processing, capable of building a scalable and high-performance solution for this project.
Related categories:
PHP
Python
WordPress
Web Scraping
MySQL
HTML
Backend Development
API Development