Node.js Expert for Parsing Data
Budget: $10 – $3,000 USD
Job Title: Expert Node.js Developer for PDF-to-CSV Parsing (Regex & AI)
Job Description
We are looking for an experienced Node.js developer to extract structured data from horse competition reports in PDF format and convert them into CSV files. The challenge is that the table structures are inconsistent, column headers vary, and AI-based parsing (GPT-4o, Gemini 2.5) has proven inaccurate.
Project Requirements
Extract summary data from the header of the report.
Identify multiple competitions within a single report.
Extract competition summaries and section-wise tables.
Handle tables with inconsistent column structures.
Ensure high accuracy (no missing or incorrect cell data).
Convert extracted data into structured CSV format.
Challenges & Issues
AI-based models (GPT-4o, Gemini 2.5) are unreliable – missing or incorrect cell data.
3rd-party services (CloudConvert) fail to extract summaries.
Regex-based parsing is difficult due to inconsistent table structures.
Deliverables
A Node.js script that extracts structured data from PDFs.
CSV files with accurate competition summaries and table data.
Documentation on how to run the script.
How to Apply
If you have experience with PDF parsing and structured data extraction, please apply with:
Your relevant experience (mention past projects).
Your approach to solving this problem.
Expected timeline and budget estimate.
Job Description
We are looking for an experienced Node.js developer to extract structured data from horse competition reports in PDF format and convert them into CSV files. The challenge is that the table structures are inconsistent, column headers vary, and AI-based parsing (GPT-4o, Gemini 2.5) has proven inaccurate.
Project Requirements
Extract summary data from the header of the report.
Identify multiple competitions within a single report.
Extract competition summaries and section-wise tables.
Handle tables with inconsistent column structures.
Ensure high accuracy (no missing or incorrect cell data).
Convert extracted data into structured CSV format.
Challenges & Issues
AI-based models (GPT-4o, Gemini 2.5) are unreliable – missing or incorrect cell data.
3rd-party services (CloudConvert) fail to extract summaries.
Regex-based parsing is difficult due to inconsistent table structures.
Deliverables
A Node.js script that extracts structured data from PDFs.
CSV files with accurate competition summaries and table data.
Documentation on how to run the script.
How to Apply
If you have experience with PDF parsing and structured data extraction, please apply with:
Your relevant experience (mention past projects).
Your approach to solving this problem.
Expected timeline and budget estimate.
Related categories:
Business, Accounting, Human Resources & Legal
JavaScript
Python
NoSQL Couch & Mongo
Node.js