Spanish Procurement Benchmark Dataset Creation

Job ID: 40565913

Budget: $10 – $30 USD

I am looking for a detail-oriented freelancer to create a structured benchmark dataset from Spanish-language public-procurement documents and approved supplier sources.
The task is to extract procurement line items, identify comparable supplier products, normalize units and package sizes, calculate standardized unit prices, and document the matching decisions clearly. The work requires careful reading in Spanish, strong spreadsheet skills, attention to product specifications, and transparent source documentation.

You will receive:
* 40 Spanish-language procurement documents or excerpts;
* detailed instructions for documenting uncertain matches, non-matches, and source limitations.


Main Task:
For each procurement item, you will:

1. Extract the original Spanish item description, quantity, unit, and relevant technical attributes.
2. Clean and normalize the product description.
3. Identify comparable supplier products in Peru
4. Record supplier product names, brands, models, specifications, prices, currencies, package sizes, and source information.
5. Convert prices into standardized unit prices.
6. Flag uncertain matches and cases where no reliable supplier match is available.
7. Provide a short rationale for each match or non-match.
8. Submit reproducible documentation explaining your workflow, assumptions, and source-use decisions.


Deliverables:
Please submit:
* `output.xlsx`: completed spreadsheet, one row per procurement item and candidate supplier match;
* `output.json`: machine-readable version of the same data;
* `README.md`: tools used, assumptions, match criteria, source-use rules, and reproduction steps;
* optional: `run.py`, `notebook.ipynb`, or another reproducible workflow file if you use automation.
You are free to do the project manually or by automation.

Each row must include:

* procurement item ID;
* original Spanish description;
* normalized product name;
* quantity and unit;
* supplier source;
* supplier product name;
* brand, model, and relevant specifications when available;
* price and currency;
* package size;
* standardized unit price;
* match-confidence score;
* match rationale;
* source access date or archive identifier.



Quality expectations:

The final dataset will be checked for extraction accuracy, source validity, product-match quality, correct unit conversion, reproducibility, and clear documentation.
Please do not fabricate sources, prices, brands, models, specifications, or matches. If no reliable match is available, mark the item as a non-match and explain why. Uncertain matches are acceptable when clearly flagged and justified.