Data Pipeline Developer

Job ID: 39860826

Budget: $250 – $750 USD

About the Role
We’re looking for a freelance data engineer to support a short-term analytics project (approx. 10–15 hours per week for 4–6 weeks). The work involves setting up an automated process to extract, clean, and organize information from PDF files and make it accessible through simple dashboards.

Key Responsibilities
Configure and refine OCR workflows using Tesseract or EasyOCR.
Develop reusable Python + Pandas scripts to clean, label, and validate extracted data.
Design and populate a PostgreSQL database and ensure clean data loading.
Connect the database to Metabase (or a similar open-source BI tool) for visualization.
Provide clear, step-by-step documentation of your process and code.

Required Skills
Proficiency in Python, especially Pandas and OCR libraries.
Understanding of database design and SQL (PostgreSQL preferred).
Experience with data visualization or dashboard tools (Metabase, Superset, etc.).
Ability to write clean, well-documented code and communicate clearly with non-technical stakeholders.

Engagement Details
Remote / Flexible hours
4–6 week engagement, hourly or fixed-rate depending on experience