Python Developer Needed for PDF Data Extraction
Budget: $30 – $250 USD
I am looking for an experienced Python developer to help me create a script that can be run in Google Colab. The goal is to extract specific data from two complex PDF files that contain extensive horse racing information. The extracted data should then be organized into a table format, with each row representing a horse and its details for a particular race.
Project Scope:
The code must be written in Python and compatible with Google Colab.
The script should allow me to upload two PDF files, both containing complex horse racing data.
The script should extract the following details per horse per race:
-Race number
-Horse name
-Jockey name
-Trainer name
-Time and distance data
-Specific data points, ie., Previous race results, speed ranking, ability, etc.
The extracted data should be organized into a clean table format (e.g., Pandas DataFrame) for easy analysis.
You should include comments in the code to ensure clarity for future modifications.
Requirements:
Experience in Python (preferably for PDF data extraction and manipulation).
Familiarity with libraries like PyPDF2, PDFMiner, or similar.
Experience working with data extraction and organization using Pandas.
Ability to work with Google Colab.
Strong attention to detail and accuracy.
Deadline: Please specify your estimated timeline for project completion.
Looking forward to your proposals!
Project Scope:
The code must be written in Python and compatible with Google Colab.
The script should allow me to upload two PDF files, both containing complex horse racing data.
The script should extract the following details per horse per race:
-Race number
-Horse name
-Jockey name
-Trainer name
-Time and distance data
-Specific data points, ie., Previous race results, speed ranking, ability, etc.
The extracted data should be organized into a clean table format (e.g., Pandas DataFrame) for easy analysis.
You should include comments in the code to ensure clarity for future modifications.
Requirements:
Experience in Python (preferably for PDF data extraction and manipulation).
Familiarity with libraries like PyPDF2, PDFMiner, or similar.
Experience working with data extraction and organization using Pandas.
Ability to work with Google Colab.
Strong attention to detail and accuracy.
Deadline: Please specify your estimated timeline for project completion.
Looking forward to your proposals!