Information Extraction & Structuring from External Sources

Job ID: 38104204

Budget: $3,000 – $5,000 USD

Project Description:

I need a developer to create a system that queries multiple external sources (18) through APIs, web scraping, and web page reengineering processes to extract information and structure it in the expected format. The system should generate a record of each query made, indicating whether results were obtained, no results were obtained, or if the source was not available. A document detailing all the sources of information from which data should be extracted will be provided.

Main Features:
1. Automatic Querying of External Sources: Perform queries to multiple external sources according to the provided specifications.
2. Data Structuring: Process the obtained information and structure it in the required format.
3. Query Log Generation: Create a detailed log of all queries made and their results.
4. Response Handling: Identify and record whether results were obtained, no results were obtained, or if the source was unavailable.
5. Executor Function: There should be a function that executes all the sources that were queried and returns a PDF file as in Annex 1.

Technical Requirements:

• Experience in the development of automated query systems and data processing from external sources.
• Knowledge in APIs, web scraping, and web page reengineering processes. Must use Java, NodeJs, or Python.
• Ability to structure data efficiently and accurately.
• Ability to generate detailed logs of queries and results.
• The solution must include the necessary files and documentation to deploy each connector or integrator separately as AWS Lambda functions, Docker containers, or any Serverless service in AWS or GCP.

Expected Deliverables:

• Complete development of the query and data structuring system with all required features.
• Implementation of an automated process to perform queries according to the provided specifications.
• Generation of a detailed log of all queries made and their results.
• Thorough testing to ensure the quality and accuracy of the system in different query scenarios.

Additional Details:
• A detailed document with all queries to be performed will be provided as a reference for the project development.
• The desired delivery time is approximately 8 weeks from the start of the project.
• Ability to propose efficient and scalable solutions for querying and structuring data from multiple external sources will be positively evaluated.
• Experience in handling proxies or any technique to avoid possible blocks due to the high concurrency of queries will be valued.

How to Submit Proposals:
Please include in your proposal a detailed description of your relevant experience, examples of similar previous projects, your approach to addressing this specific project, and an estimate of costs and deadlines.
Related categories: Python Web Scraping Data Mining Aws Lambda ETL