PDF Content Extraction Script Development

Job ID: 37559705

Budget: $250 – $750 USD

I'm in need of a skilled developer to create an automated script to accurately scrape text content from large-sized PDF files, ranging up to 100TB.

Key Project Requirements:
* Development of a script able to extract textual content from PDF files effectively.
* Ability to handle PDF files of large size, up to 100TB.
* Deal with content that's simple text only, no need to process images or tables embedded in the PDF.

I am looking for a freelancer who can help me automate the search in a large number of PDFs and categorize the output keyword information for data analysis purposes.

Requirements:
- Experience in scripting and automation
- Strong knowledge of PDF processing
- Proficiency in extracting and categorizing keywords
- Ability to work with large datasets 100TB

The project involves processing approximately 120 million PDFs and extracting keywords from them. The extracted information should be organized and categorized into a CSV file for further analysis.

Ideal Skills and Experience:
* Familiarity with PDF content scraping techniques.
* Proficient in a scripting language suitable for this task.
* Experience in handling and processing large datasets.
* Strong problem-solving skills to manage potential challenges posed by the substantial size of the files.
Related categories: Python Data Entry Excel Web Scraping Data Mining