Full-Stack Development for Data Acquisition Platform

Job ID: 40514932

Budget: ₹12,500 – ₹37,500 INR

Enterprise-Grade Data Acquisition & Document Intelligence Platform

Project Description:

We are seeking an experienced Full-Stack Developer or Development Team to build the MVP of a large-scale document acquisition, processing, and intelligence platform.

The objective is to automatically collect documents and metadata from multiple public and semi-public web sources, normalize the data into a unified structure, store documents in a central repository, and expose the information through APIs and administrative dashboards.

Scope of Work:

• Build scalable web crawlers and source-specific adapters.
• Automate document discovery and downloading (PDF, DOC, DOCX, XLS, ZIP and other formats).
• Extract and normalize metadata from multiple sources into a unified schema.
• Implement OCR pipeline for scanned documents.
• Store structured data and downloaded documents in a centralized database and file repository.
• Build scheduling, monitoring, retry and failure-handling mechanisms.
• Develop REST APIs for data access and integration.
• Create an administrative dashboard for monitoring crawler health, data quality and processing status.
• Implement logging, audit trails and operational reporting.
• Containerize the complete solution using Docker.
• Provide deployment documentation and technical handover.

Technical Requirements:

• Python preferred for crawling and data processing.
• PostgreSQL (or equivalent enterprise-grade database).
• Docker-based deployment.
• REST API architecture.
• OCR integration.
• Document parsing and metadata extraction.
• Scalable architecture capable of supporting future expansion.

Deliverables:

• Complete source code.
• Database schema.
• Docker configuration.
• Deployment documentation.
• API documentation.
• Installation guide.
• Administrator guide.
• Full ownership and transfer of intellectual property.

Important Conditions:

• The solution must be developer-independent and fully portable.
• No vendor lock-in.
• All source code and documentation must be delivered.
• The system must be maintainable by future development teams.
• The architecture should support future AI, analytics and intelligence modules without requiring major redesign.

Required Experience:

• Large-scale web scraping and crawling.
• OCR and document processing.
• Data engineering.
• Backend/API development.
• PostgreSQL and database optimization.
• Docker and deployment automation.
• Enterprise software architecture.

Please include relevant project examples, proposed technology stack, estimated timeline, team composition, and fixed-price quotation.