Procurement Data Extraction from Spanish Documents -- 2
Budget: $10 – $30 USD
We need a freelancer to: 1. extract target procurement line items from Spanish documents; The number of documents are 20 pdfs. 2. clean and normalize descriptions, quantities, units, brands, models, and technical attributes; 3. identify comparable supplier products using the approved source environment; 4. extract prices, currencies, package sizes, and source information; 5. convert prices into standardized unit prices; 6. flag uncertain matches and cases where no reliable match is available; 7. submit a reproducible data file and workflow documentation.
Deliverables are:
• output.xlsx: one row per procurement item and candidate supplier match;
• output.json: a machine-readable version of the same data;
• run.py, notebook.ipynb, or equivalent reproducible workflow, when automation is used. (You are free to do it manually or by automation);
• README.md: tools, assumptions, source-use rules, match criteria, and reproduction steps.
Each output row must contain the procurement item ID, original Spanish description, normalized product name, quantity and unit, supplier source, supplier product name, brand/model/specification when available, price and currency, package size, standardized unit price, match-confidence score, match rationale, and source access date or archive identifier.
Deliverables are:
• output.xlsx: one row per procurement item and candidate supplier match;
• output.json: a machine-readable version of the same data;
• run.py, notebook.ipynb, or equivalent reproducible workflow, when automation is used. (You are free to do it manually or by automation);
• README.md: tools, assumptions, source-use rules, match criteria, and reproduction steps.
Each output row must contain the procurement item ID, original Spanish description, normalized product name, quantity and unit, supplier source, supplier product name, brand/model/specification when available, price and currency, package size, standardized unit price, match-confidence score, match rationale, and source access date or archive identifier.
Related categories:
Python
Data Processing
Data Entry
Excel
Data Scraping
Data Extraction
Data Analysis
Data Integration
Data Management
Pandas