Junior AI Post-Training Analyst

Job ID: 40570047

Budget: $15 – $25 USD

My language-model experiments have finished training and I now need a junior engineer who can dive into the results and turn raw logs into clear, actionable insights. All work centers on text data that has already been pre-cleaned and trained; your focus is firmly on the post-training stage.

You will
• parse evaluation logs, confusion matrices, and attention weights coming from Python pipelines,
• calculate and visualise key metrics (accuracy, F1, perplexity, BLEU, etc.),
• compare successive model checkpoints,
• write concise reports that highlight strengths, failure cases, and concrete next-step recommendations,
• automate the above in repeatable Python or Java scripts so future runs can be analysed just as easily.

Required know-how: solid Python (pandas, NumPy, matplotlib or seaborn; Jupyter is fine) plus enough Java to work with existing evaluation utilities. Familiarity with scikit-learn or Hugging Face evaluation tools will help you get up to speed quickly.

Deliverables
1. A reproducible notebook or script that ingests the model’s JSON/CSV output and produces metric tables and plots.
2. A short markdown/PDF report summarising findings and recommending parameter tweaks or data augmentation ideas.
3. Clean, well-commented code pushed to our private repo and a brief hand-off call or video walkthrough.

Acceptance criteria: scripts must run end-to-end on our sample dataset without manual edits, and the report should clearly flag at least three performance bottlenecks with evidence-backed suggestions.

If this feels like the right next step in your AI journey, let’s get started—I’m ready to share the evaluation data as soon as we agree on the timeline.