Design the Architecture of Pharmelis Registry (Global Pharmaceutical Data System)
Budget: €36 – €0 EUR
We are building Pharmelis Registry — a canonical database for pharmaceuticals.
To make any pharmaceutical product understandable, anywhere, in any language.
Pharmaceutical data today is fragmented, inconsistent, and multilingual. There is no reliable way to identify and reconcile products globally across sources and markets.
Pharmelis Registry aims to capture real-world data and resolve it into clear, unambiguous product identities.
Objective
Design a system that:
- ingests heterogeneous pharmaceutical data (CSV, APIs, websites, PDFs, images)
- works across countries and languages
- handles messy, inconsistent data
- resolves product identity across sources
- produces a consistent global reference
- includes an internal dashboard to operate the system
The dashboard must allow:
- inspecting ingested data
- reviewing identity decisions
- monitoring system coverage
- adding and managing data sources
Deliverables
Provide a concise architecture document covering:
- identity resolution strategy (core part)
- ingestion approach for different source types
- core data model (what is stored vs computed)
- system architecture and data flow
- dashboard design and capabilities
- recommended tech stack and trade-offs
Required Profile
Strong in:
- Python
- data engineering / pipelines
- web scraping / ingestion
- API/backend development
- dashboard or web app development
Experience with messy data, search systems, or entity resolution is expected.
To make any pharmaceutical product understandable, anywhere, in any language.
Pharmaceutical data today is fragmented, inconsistent, and multilingual. There is no reliable way to identify and reconcile products globally across sources and markets.
Pharmelis Registry aims to capture real-world data and resolve it into clear, unambiguous product identities.
Objective
Design a system that:
- ingests heterogeneous pharmaceutical data (CSV, APIs, websites, PDFs, images)
- works across countries and languages
- handles messy, inconsistent data
- resolves product identity across sources
- produces a consistent global reference
- includes an internal dashboard to operate the system
The dashboard must allow:
- inspecting ingested data
- reviewing identity decisions
- monitoring system coverage
- adding and managing data sources
Deliverables
Provide a concise architecture document covering:
- identity resolution strategy (core part)
- ingestion approach for different source types
- core data model (what is stored vs computed)
- system architecture and data flow
- dashboard design and capabilities
- recommended tech stack and trade-offs
Required Profile
Strong in:
- Python
- data engineering / pipelines
- web scraping / ingestion
- API/backend development
- dashboard or web app development
Experience with messy data, search systems, or entity resolution is expected.