AI Medical PDF Extraction
Budget: €30 – €250 EUR
I’m looking for a reliable way to turn a handful of text-based PDF directories of medical professionals into structured data I can use immediately. Each record in the PDFs shows a practitioner’s full name followed by their qualification, and I need four clean columns in Google Sheets or Excel:
• Name
• Surname
• Qualification
• Specialty (mapped from my existing reference list)
Because many entries include multiple given names, initials or suffixes, a simple split won’t be accurate enough. Please build an AI-driven workflow—whether that’s Python with NLP libraries, a no-code AI extractor, or another proven method—that:
1. Reads the text-based PDFs directly (no OCR required).
2. Separates first name(s) and surname with high accuracy.
3. Captures the qualification exactly as written.
4. Looks up that qualification against my supplied list and inserts the matching specialty.
5. Outputs a ready-to-use Google Sheet or Excel file and the reusable script/notebook so I can run it on new PDFs later.
If you have experience with large-scale text parsing, entity recognition, or data alignment for healthcare datasets, this should be straightforward. Accuracy in name splitting and specialty mapping is the priority; I’ll verify by spot-checking random rows. Let me know how soon you can deliver the first pass and any libraries or services you plan to rely on.
• Name
• Surname
• Qualification
• Specialty (mapped from my existing reference list)
Because many entries include multiple given names, initials or suffixes, a simple split won’t be accurate enough. Please build an AI-driven workflow—whether that’s Python with NLP libraries, a no-code AI extractor, or another proven method—that:
1. Reads the text-based PDFs directly (no OCR required).
2. Separates first name(s) and surname with high accuracy.
3. Captures the qualification exactly as written.
4. Looks up that qualification against my supplied list and inserts the matching specialty.
5. Outputs a ready-to-use Google Sheet or Excel file and the reusable script/notebook so I can run it on new PDFs later.
If you have experience with large-scale text parsing, entity recognition, or data alignment for healthcare datasets, this should be straightforward. Accuracy in name splitting and specialty mapping is the priority; I’ll verify by spot-checking random rows. Let me know how soon you can deliver the first pass and any libraries or services you plan to rely on.