Scientific Python Engineer – GC-MS Analytical Modeling & Constrained Optimization

Job ID: 40239436

Budget: $750 – $1,500 USD

Project Overview

We are developing an internal AI-driven GC-MS analytical modeling system for flavor formulation reverse engineering.

This is not:

A chatbot project

A basic automation workflow

A generic machine learning classification task

A web dashboard build

This is a scientific inverse modeling and constrained optimization system built around GC-MS analytical data.

If you do not have experience with scientific modeling, regression, and constrained optimization, please do not apply.

System Objectives

We are building a modular Python system that:

Parses structured GC-MS peak lists (CSV/Excel, sometimes PDF reports).

Performs CAS-based compound normalization and matching.

Handles retention time alignment.

Applies weighted similarity scoring with marker compound prioritization.

Uses paired historical datasets (Recipe % ↔ GC-MS area %) to statistically estimate response factors.

Generates optimized formulation trials under strict constraints:

Sum = 100%

Ingredient min/max limits

Regularization to prevent micro-dosing overfitting

Performs residual minimization after re-analysis.

Supports iterative refinement (Trial-1 → Trial-2 → Trial-3 loop).

This is an inverse analytical modeling engine, not a surface-level similarity system.

Mandatory Requirements (No Exceptions)

You must have:

Strong Python experience (minimum 4–5 years)

Advanced Pandas and NumPy knowledge

Experience with scientific or analytical datasets

Experience implementing regression models

Experience with constrained optimization (linear or nonlinear)

Understanding of L1/L2 regularization

Experience with numerical modeling (SciPy optimize or similar)

Ability to clearly explain statistical calibration logic

Clean modular code architecture skills

If you cannot explain:

Weighted least squares regression

Constrained nonlinear optimization

Regularization to prevent overfitting

Residual minimization logic

Please do not apply.

Preferred Experience

GC-MS, LC-MS, chromatography, or spectral data processing

Peak alignment techniques

NIST or compound library matching

Background in analytical chemistry, engineering, or physics

Experience building internal scientific R&D tools

Project Structure

Phase 1:

GC-MS parser

CAS normalization

Peak alignment

Structured dataset engine

Phase 2:

Marker-aware similarity scoring

Statistical response factor estimation (regression-based)

Phase 3:

Constrained optimization engine

Regularization and residual minimization

Iterative trial refinement framework

Application Requirements

To apply, you must answer all of the following:

How would you estimate response factors between ingredient % usage and GC-MS area % using paired historical data?

What optimization method would you use to generate a 100% normalized recipe under ingredient constraints?

How would you prevent overfitting to minor peaks?

Have you worked with chromatographic or mass spectrometry data before? Provide details.

Applications without technical answers will not be considered.

This is a long-term internal scientific R&D project. If your experience is primarily in chatbots, generic AI automation, or web SaaS tools, this project is likely not a match.