Data Engineer — Insurance Claims Data Pipeline (ETL + Data Quality + Light AI)
Budget: $30 – $250 USD
We need a data engineer for a small, self-contained project building a data pipeline for insurance claims data. Most of the finer details will be worked out over chat, so this is just to give you the shape of it.
The gist: we have claims data coming from a few different sources in inconsistent, messy formats — structured exports plus some free-text notes. We want a pipeline that ingests it, cleans and validates it (handling duplicates, conflicting records, bad/missing values properly rather than silently dropping them), lands it in a clean, well-modeled schema, and produces a simple prioritized "review this first" output plus a few summary metrics.
There's also a light AI/ML component — for example, using an LLM to pull structured fields out of the free-text notes, or a simple model to score/rank claims for review. Nothing exotic; we care more about it being done sensibly and reliably than about model complexity.
What we're looking for:
- Strong SQL and hands-on ETL / data engineering experience
- Real data-quality and validation instincts (this is the core of the job)
- Comfort with Python and basic AI/LLM or ML usage
* Bonus: experience with insurance, fintech, or other regulated-industry data
Scope is roughly a few days of work. The dataset is synthetic — no real personal data — and will be provided. Please mention relevant experience (especially any claims/insurance/financial data work) in your bid.
The gist: we have claims data coming from a few different sources in inconsistent, messy formats — structured exports plus some free-text notes. We want a pipeline that ingests it, cleans and validates it (handling duplicates, conflicting records, bad/missing values properly rather than silently dropping them), lands it in a clean, well-modeled schema, and produces a simple prioritized "review this first" output plus a few summary metrics.
There's also a light AI/ML component — for example, using an LLM to pull structured fields out of the free-text notes, or a simple model to score/rank claims for review. Nothing exotic; we care more about it being done sensibly and reliably than about model complexity.
What we're looking for:
- Strong SQL and hands-on ETL / data engineering experience
- Real data-quality and validation instincts (this is the core of the job)
- Comfort with Python and basic AI/LLM or ML usage
* Bonus: experience with insurance, fintech, or other regulated-industry data
Scope is roughly a few days of work. The dataset is synthetic — no real personal data — and will be provided. Please mention relevant experience (especially any claims/insurance/financial data work) in your bid.
Related categories:
Python
Data Processing
SQL
Insurance
Paralegal Services
Data Analysis
ETL
FinTech
Data Engineer
LLM Integration