AI-Driven GC-MS Automation Development

Job ID: 40239223

Budget: $750 – $1,500 USD

Python Developer – GC-MS Data Analysis & Automated Flavor Recipe Generation System

Project Overview

We are a flavor manufacturing company developing an internal AI-driven GC-MS automation system.

The goal of this project is to:

Process GC-MS peak data

Match peaks with compound libraries

Identify possible natural/synthetic sources

Detect marker compounds

Estimate ingredient combinations

Automatically generate optimized trial recipes (normalized to 100%)

This is NOT a simple chatbot or automation project.
This is a scientific data processing and algorithm development project.

Responsibilities

Parse GC-MS peak lists (CSV / Excel exports)

Implement compound matching logic (CAS-based matching)

Normalize peak area percentages

Develop similarity scoring algorithms

Identify marker compounds and decision rules

Create probabilistic ingredient combination models

Generate optimized recipe outputs (percentage-based, normalized to 100%)

Structure the system in a modular and scalable Python architecture

Build clean, documented code suitable for long-term expansion

Required Skills (Mandatory)

Strong Python programming experience (3+ years)

Advanced Pandas & NumPy knowledge

Scientific data processing experience

Experience working with structured chemical datasets

Algorithm development experience

Data normalization & similarity scoring logic

CSV / Excel data parsing

Clean code & modular architecture mindset

Strong Plus

Experience with chromatography or GC-MS data

Background in analytical chemistry

Experience in scientific computing

Experience building internal AI decision systems

Experience with LLM integration for explanation layers

What This Is NOT

Not a chatbot-only project

Not a no-code automation task

Not a simple API integration

Not a website project

This is a data science & algorithm-driven system.

Deliverables

Phase 1:

GC-MS parser module

Compound matching engine

Similarity scoring model

Phase 2:

Ingredient probability engine

Marker compound logic

Recipe generation module

Phase 3:

Optimization & refinement engine

Modular expansion framework

To Apply, Please Answer:

Have you worked with chromatography or GC-MS data before?

Explain how you would normalize peak area percentages.

How would you design a similarity scoring algorithm for compound matching?

Share an example of a scientific data processing project you built.

Which Python libraries would you use for this project and why?